Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI Apps

Why AI Features Get Worse, and How to Notice

Last updated:

Conventional software does not do this

Ordinary software keeps doing exactly what it did. An AI feature that was 94% accurate in March can be materially worse in September with nothing changed on your side.

That is the single biggest operational difference, and it is why AI maintenance costs more than conventional maintenance.

Three causes of drift

  1. Your inputs change. New document formats, new customer types, new product lines, seasonal language.
  2. Your ground truth changes. A policy is updated, so answers that were right are now wrong.
  3. The model changes. Providers update models, sometimes with subtle behavioural differences.

What we monitor

  • Correction rate by category — the earliest reliable signal
  • Refusal rate — a jump usually means retrieval has broken rather than the model
  • Output length distribution, which shifts before quality visibly does
  • Confidence distribution — more items landing in the review band
  • Scheduled evaluation runs against the fixed test set
Uptime dashboards give false comfort with AI. A model that is up and answering badly is worse than one that is down, because nothing alerts.

Refresh the evaluation set

A test set built in January describes January. Add new real cases periodically, especially ones the system got wrong, and retire cases that no longer represent your work.

Hold back a portion you never tune against, or you will eventually optimise for the test rather than for reality.

Budget for it

Twenty to thirty per cent of build cost annually, higher than conventional software. That covers evaluation reruns, threshold tuning, prompt adjustments after provider changes, and refreshing the test set.

A business that budgets nothing ends up with a feature nobody trusts, which is worse than not having built it.

Frequently asked questions

How often should we rerun evaluation?

Automatically on every change, and on a schedule — weekly is reasonable for anything customer-facing. It should not depend on anyone remembering.

Can we pin the model version?

Where the provider allows it, yes, and it delays rather than removes the problem since versions are eventually retired. Pin, then move deliberately with the test set.

Who owns quality after launch?

Someone in the business who cares about the outcome, supported by whoever can change the system. Unowned AI features degrade invisibly.

What if quality drops and we cannot fix it?

Widen the review band temporarily so more goes to humans, then diagnose. Degrading gracefully beats running wrong.

Keep reading

Have an AI feature nobody is measuring?

It may already have drifted. We can build an evaluation harness around a system we did not write.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsCustom Software DevelopmentSaaS Development