Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. How to Measure Whether an AI Project Actually Worked
AI & Machine Learning

How to Measure Whether an AI Project Actually Worked

A measurement approach for AI projects that survives scrutiny, including the baseline most teams forget to take.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Take a baseline before you build, measure the same things after, and count only benefits you can point at. Time saved counts if the time was redeployed; error reduction counts if you can price the errors; capacity counts if you used it. Everything else is narrative.

The baseline is the whole game

Most AI projects cannot prove their value because nobody measured the before. Six months later the debate is between people who feel it helped and people who feel it did not, and neither can settle it.

Two weeks of measurement before anything changes converts that argument into arithmetic. It is the cheapest insurance available on an AI budget.

What to measure before you start

  1. Time per unit of work — minutes per ticket, per document, per enquiry.
  2. Volume handled per week.
  3. Error or rework rate, however roughly you can capture it.
  4. Elapsed time from start to finish, which is different from effort.
  5. The cost of the current process, including the people doing it.

Measure by observation where you can. Self-reported figures are systematically wrong and will be challenged later by whoever does not like the conclusion.

Count benefits you can actually point at

  • Time redeployed — only if you can say what the hours went to instead
  • Errors avoided, priced from real historical error costs
  • Volume handled without hiring, which is real capacity
  • Revenue attributable to faster response, where you can show the link
  • Reduced supplier or licence costs that actually stopped being paid
The discipline: if you cannot name the invoice that stopped or the hire that did not happen, it is a soft benefit. Soft benefits are real but they belong in a separate section of the paper, clearly labelled.

Count the full cost

Build cost, model and infrastructure running costs, the human review that continues, internal time on specification and testing, and ongoing maintenance and evaluation.

AI features have a higher ongoing cost than conventional software because they need monitoring and periodic re-evaluation. A business case that treats the build as the whole cost will look wrong within a year.

Attribution honestly

If you changed the process at the same time as introducing AI — and you almost certainly did — some of the benefit belongs to the process change. Where possible, phase them: change the process first, measure, then add the AI.

Where phasing is impractical, say so in the paper rather than claiming the whole benefit. Overclaiming on the first project makes the second one harder to fund when someone checks.

Review on a schedule and be willing to stop

Set a review at three months and six months with the same metrics. Decide in advance what result would mean widening, iterating or stopping.

Stopping a project that did not work is a good outcome, not a failure. The failure is a system that quietly persists for two years because nobody wanted to be the one to say it was not helping.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

What if we did not take a baseline and have already built it?

Reconstruct what you can from historical records — ticket volumes, handling times, error logs — and be explicit about the uncertainty. Then take a proper baseline before the next project.

How long before we should expect a return?

For process automation with AI components, six to eighteen months is a reasonable expectation. Anything promising a return in weeks is usually counting benefits that have not been banked.

Should we count improved quality?

Yes, if you can measure it — error rates, satisfaction scores, rework. “Better quality” without a measurement is a claim rather than a benefit.

Who should own the measurement?

Someone other than the person who championed the project, ideally. Not from distrust, but because independent measurement is more credible when the results are presented.

Keep reading

More on AI & Machine Learning

Start here

About to start an AI project?

Take the baseline first. We will tell you which five numbers to capture and how, whoever ends up building the thing.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →