Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Why AI Features Get Worse, and How to Notice
AI Apps

Why AI Features Get Worse, and How to Notice

Why AI apps lose accuracy over time with no code changes: the three causes of drift, what to monitor, and refreshing the test set to keep quality steady.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

AI features degrade without any code change: your documents change, your customers change, the provider updates the model. Automated evaluation reruns, drift alerts on refusal and correction rates, and periodic refresh of the test set are what keep quality steady.

Conventional software does not do this

Ordinary software keeps doing exactly what it did. An AI feature that was 94% accurate in March can be materially worse in September with nothing changed on your side.

That is the single biggest operational difference, and it is why AI maintenance costs more than conventional maintenance.

Three causes of drift

  1. Your inputs change. New document formats, new customer types, new product lines, seasonal language.
  2. Your ground truth changes. A policy is updated, so answers that were right are now wrong.
  3. The model changes. Providers update models, sometimes with subtle behavioural differences.

What we monitor

  • Correction rate by category — the earliest reliable signal
  • Refusal rate — a jump usually means retrieval has broken rather than the model
  • Output length distribution, which shifts before quality visibly does
  • Confidence distribution — more items landing in the review band
  • Scheduled evaluation runs against the fixed test set
Uptime dashboards give false comfort with AI. A model that is up and answering badly is worse than one that is down, because nothing alerts.

Refresh the evaluation set

A test set built in January describes January. Add new real cases periodically, especially ones the system got wrong, and retire cases that no longer represent your work.

Hold back a portion you never tune against, or you will eventually optimise for the test rather than for reality.

Budget for it

Twenty to thirty per cent of build cost annually, higher than conventional software. That covers evaluation reruns, threshold tuning, prompt adjustments after provider changes, and refreshing the test set.

A business that budgets nothing ends up with a feature nobody trusts, which is worse than not having built it.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

How often should we rerun evaluation?

Automatically on every change, and on a schedule — weekly is reasonable for anything customer-facing. It should not depend on anyone remembering.

Can we pin the model version?

Where the provider allows it, yes, and it delays rather than removes the problem since versions are eventually retired. Pin, then move deliberately with the test set.

Who owns quality after launch?

Someone in the business who cares about the outcome, supported by whoever can change the system. Unowned AI features degrade invisibly.

What if quality drops and we cannot fix it?

Widen the review band temporarily so more goes to humans, then diagnose. Degrading gracefully beats running wrong.

Keep reading

More on AI Apps

AI Apps

The Hidden Work in 'Simple' AI Features

Why an AI feature that took an afternoon to demo takes weeks to ship: the evaluation, edge cases, guardrails, cost control and monitoring nobody sees.

Start here

Have an AI feature nobody is measuring?

It may already have drifted. We can build an evaluation harness around a system we did not write.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →