Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI Apps

Designing a Pilot You Can Draw a Conclusion From

Last updated:

Most pilots succeed and change nothing

The characteristic failure is not a pilot that fails. It is one that goes reasonably well, everyone is pleased, and eighteen months later it is still a pilot because nobody agreed what would justify going further.

The fix is deciding the threshold in advance, while nobody is invested in the answer.

Four things to agree before it starts

  1. The measure — accuracy on what, straight-through rate, time per item
  2. The threshold — the number that means proceed
  3. The sample — how many cases, chosen how, representative of what
  4. The duration and who decides at the end
Write the threshold down and circulate it. The version of this conversation held after the results are in is considerably less productive.

Make the sample representative

Pilots run on convenient data produce convenient results. Include the awkward cases, the seasonal variation and the customer types you find difficult, in roughly the proportion they actually occur.

A pilot on your ten best-behaved suppliers tells you nothing about your supplier base.

Shadow rather than parallel where you can

Have the system produce output that nobody acts on while humans do the work as usual, then compare. That gives you a real accuracy measurement with no operational risk at all.

It also avoids the confound where people change their behaviour because they know the pilot is running.

Decide, in writing, at the end

  • Above threshold: proceed to build, with the scope defined
  • Close: one specific improvement, retested, with a date
  • Below: stop, and write down why

Stopping is a legitimate outcome and it is far better than a pilot that persists indefinitely because nobody wants to conclude it.

Frequently asked questions

How long should a pilot run?

Long enough to cover a full cycle including any monthly peak — usually four to six weeks. Shorter samples miss the variation that matters.

Who should decide the threshold?

The person who will own the outcome, informed by what the current process achieves. “Better than we do now, at acceptable risk” is a defensible starting point.

What if the result is ambiguous?

That usually means the sample was too small or the measure was wrong. Both are fixable, and both are better than proceeding on an ambiguous result.

Should the pilot use the real interface?

Only enough to see the output. Building a polished interface for a pilot spends money on the wrong question.

Keep reading

About to run an AI pilot?

Agree the threshold before it starts, whoever runs it. Happy to help you set one that means something.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsCustom Software DevelopmentSaaS Development