How to Measure Whether an AI Project Actually Worked
Last updated:
The baseline is the whole game
Most AI projects cannot prove their value because nobody measured the before. Six months later the debate is between people who feel it helped and people who feel it did not, and neither can settle it.
Two weeks of measurement before anything changes converts that argument into arithmetic. It is the cheapest insurance available on an AI budget.
What to measure before you start
- Time per unit of work — minutes per ticket, per document, per enquiry.
- Volume handled per week.
- Error or rework rate, however roughly you can capture it.
- Elapsed time from start to finish, which is different from effort.
- The cost of the current process, including the people doing it.
Measure by observation where you can. Self-reported figures are systematically wrong and will be challenged later by whoever does not like the conclusion.
Count benefits you can actually point at
- Time redeployed — only if you can say what the hours went to instead
- Errors avoided, priced from real historical error costs
- Volume handled without hiring, which is real capacity
- Revenue attributable to faster response, where you can show the link
- Reduced supplier or licence costs that actually stopped being paid
The discipline: if you cannot name the invoice that stopped or the hire that did not happen, it is a soft benefit. Soft benefits are real but they belong in a separate section of the paper, clearly labelled.
Count the full cost
Build cost, model and infrastructure running costs, the human review that continues, internal time on specification and testing, and ongoing maintenance and evaluation.
AI features have a higher ongoing cost than conventional software because they need monitoring and periodic re-evaluation. A business case that treats the build as the whole cost will look wrong within a year.
Attribution honestly
If you changed the process at the same time as introducing AI — and you almost certainly did — some of the benefit belongs to the process change. Where possible, phase them: change the process first, measure, then add the AI.
Where phasing is impractical, say so in the paper rather than claiming the whole benefit. Overclaiming on the first project makes the second one harder to fund when someone checks.
Review on a schedule and be willing to stop
Set a review at three months and six months with the same metrics. Decide in advance what result would mean widening, iterating or stopping.
Stopping a project that did not work is a good outcome, not a failure. The failure is a system that quietly persists for two years because nobody wanted to be the one to say it was not helping.
Frequently asked questions
What if we did not take a baseline and have already built it?
How long before we should expect a return?
Should we count improved quality?
Who should own the measurement?
About to start an AI project?
Take the baseline first. We will tell you which five numbers to capture and how, whoever ends up building the thing.