How to Roll Out AI Without a Big Bang
Last updated:
Three phases, each with a gate
| Phase | What happens | Gate to the next |
|---|---|---|
| Shadow | AI runs, output logged, nobody acts | Accuracy above the agreed threshold |
| Assist | Suggestions shown, people decide | Acceptance rate stable for four weeks |
| Automate | Confident cases proceed alone | Error rate below the manual baseline |
Skipping straight to the third phase is how integrations get switched off in week two.
Shadow is the cheapest insurance you will buy
Two weeks of the system producing output nobody acts on gives you a real accuracy figure at zero operational risk. It is the single most useful thing you can do before launching.
It also tells you which categories are weak, so the thresholds are set from evidence rather than from optimism.
Assist is where the value starts
Most of the time saving arrives in phase two, not phase three. Suggestions that a person accepts quickly are nearly as fast as full automation, with all the safety of human judgement.
Plenty of good integrations stop here permanently, and that is a legitimate destination rather than a failure.
Automate narrowly
- Only the categories with proven high accuracy
- Only above a confidence threshold set from real data
- With sampling — a proportion still reviewed, forever
- With a switch that returns everything to assist mode instantly
Communicating each phase
Tell the team what phase you are in and what it means for them. Ambiguity about whether the machine or the person is responsible is the fastest route to something being missed.
One paragraph per phase, circulated in advance, prevents most of the confusion.
Frequently asked questions
How long is each phase?
Can we skip shadow if we are confident?
What if acceptance rate never stabilises?
Do we have to reach phase three?
Nervous about letting AI touch a live process?
Start in shadow. Tell us the process and we will design the phases with the gates written down.