Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Measuring Whether Your AI Feature Is Any Good
AI & Machine Learning

Measuring Whether Your AI Feature Is Any Good

Building an evaluation set, choosing metrics that reflect the business need, and avoiding the traps in AI measurement.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Build a set of real cases with agreed correct answers, decide what counts as correct before you measure, and track the metric that reflects the business consequence. Without this you are tuning on anecdote, which reliably makes systems worse over time.

Anecdote is the default and it is terrible

Without a measurement set, AI systems are improved by whoever complained most recently. Someone reports a bad answer, the prompt is adjusted, something else degrades, and nobody notices because nobody is measuring.

The fix is unglamorous: a fixed set of cases you re-run on every change.

Building the set

  1. Take real cases, not invented ones. Real inputs are messier and that is the point.
  2. Include the awkward ones deliberately — the ambiguous, the incomplete, the unusual.
  3. Agree the correct answer with someone who knows the domain, in writing.
  4. Include cases where the right answer is “I do not know”, because refusal behaviour needs testing too.
  5. Aim for a few hundred, which is enough for most business features.
The disagreements that arise while agreeing correct answers are themselves valuable. If two experts disagree on 20% of cases, no system will exceed 80% agreement and you have learned the ceiling before building.

Choose metrics that reflect consequence

Overall accuracy hides the failures that matter. A system that is 95% accurate but wrong on your highest-value cases is worse than one that is 90% accurate uniformly.

  • Measure by category, not just overall
  • Weight by business consequence where errors differ in cost
  • Track false positives and false negatives separately — they usually cost different amounts
  • Report straight-through rate, not just field accuracy, where human review is involved

Use a model to grade, carefully

For subjective outputs, having a model evaluate against criteria is a practical approach and needs validating: check its grades against human grades on a sample before trusting them.

Where the criteria are objective, prefer deterministic checks. A model grading whether a date was extracted correctly is an unnecessary source of noise.

Watch for the trap of overfitting to the set

If you tune repeatedly against the same cases, you will eventually optimise for them rather than for the real distribution. Hold back a portion you do not tune against, and refresh the set periodically with new real cases.

Also re-examine the set when the world changes: new document formats, new product lines, new customer types.

Run it automatically

The evaluation should run on every change, without anyone remembering. Manual evaluation happens twice and then stops, at which point you are back to anecdote.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

How long does building an evaluation set take?

Typically one to two weeks including the domain expert's time to agree correct answers. It is the best-value fortnight in most AI projects.

What accuracy is good enough?

Whatever beats the current process at acceptable risk, with a good escalation path. Decide the threshold before you are emotionally invested in shipping.

Can we use production data as our evaluation set?

Use it as the source, and you need agreed correct answers, which production data does not come with. Sampling production and labelling it is the usual approach.

Who should own evaluation?

Someone in the business who understands what correct means, working with whoever builds it. Evaluation owned entirely by the builder tends to measure what is easy rather than what matters.

Keep reading

More on AI & Machine Learning

Start here

Improving an AI feature by guesswork?

An evaluation set turns that into measurement. Tell us what your feature does and we will suggest what the set should contain.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →