Most pilots succeed and change nothing
The characteristic failure is not a pilot that fails. It is one that goes reasonably well, everyone is pleased, and eighteen months later it is still a pilot because nobody agreed what would justify going further.
The fix is deciding the threshold in advance, while nobody is invested in the answer.
Four things to agree before it starts
- The measure — accuracy on what, straight-through rate, time per item
- The threshold — the number that means proceed
- The sample — how many cases, chosen how, representative of what
- The duration and who decides at the end
Write the threshold down and circulate it. The version of this conversation held after the results are in is considerably less productive.
Make the sample representative
Pilots run on convenient data produce convenient results. Include the awkward cases, the seasonal variation and the customer types you find difficult, in roughly the proportion they actually occur.
A pilot on your ten best-behaved suppliers tells you nothing about your supplier base.
Shadow rather than parallel where you can
Have the system produce output that nobody acts on while humans do the work as usual, then compare. That gives you a real accuracy measurement with no operational risk at all.
It also avoids the confound where people change their behaviour because they know the pilot is running.
Decide, in writing, at the end
- Above threshold: proceed to build, with the scope defined
- Close: one specific improvement, retested, with a date
- Below: stop, and write down why
Stopping is a legitimate outcome and it is far better than a pilot that persists indefinitely because nobody wants to conclude it.