Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
Software Strategy

Signs Your AI Project Is Heading for Trouble

Last updated:

Projects do not fail suddenly

AI projects rarely end with a dramatic failure. They drift. Updates get vaguer, the go-live date moves by a month, then another, and eventually the budget holder stops asking. By then most of the money is gone.

The signs are there much earlier. Some are general project warnings, which we covered in signs your software project is in trouble. The ones below are specific to machine learning and AI work, where the uncertainty is higher and the problems hide more easily behind technical language.

Warning signs in how progress is reported

  1. The success measure has changed more than once. It started as 'reduce manual review by half', became 'improve accuracy', and is now 'build a foundation for future AI'. Each change moves the goalposts closer to wherever the team happens to be.
  2. Every number is a test-set score. Months in, nobody can show performance on live or recent data. Test scores are easy to make look good.
  3. Updates describe activity, not results. 'We tried three new architectures' is activity. 'Error on last month's invoices fell from 12% to 8%' is a result.
  4. Accuracy keeps improving by tiny amounts. Weeks spent moving from 88.1% to 88.6% usually means the team is tuning because the real blockers, such as data or integration, are harder to face.

Warning signs in the data

  1. Data issues are always 'nearly sorted'. If access to a key source has been a week away for two months, something political or technical is wrong and needs escalating.
  2. Labels are being argued about. Two experts disagree about what the right answer is for a quarter of examples. The model cannot learn a distinction the business has not agreed on.
  3. Results look too good. A sudden jump to 98% is more often data leakage than a breakthrough. Treat surprise good news with the same suspicion as bad.
  4. The training data is old or hand-picked and nobody has tested on what arrived last month.

Warning signs in people and process

  1. No end user has touched it. If the people who will rely on the model have not seen a real prediction by the halfway point, adoption will be the next crisis.
  2. The decision owner has gone quiet, stopped attending reviews or delegated them to someone junior. Interest has moved on, and so will funding.
  3. Nobody has asked how it will be deployed. Integration is still described as 'the easy part at the end'. It rarely is.
  4. For language model projects, there is no evaluation set. Quality is judged by someone reading a few outputs and saying they look good. Every prompt change is a gamble.

How serious is it?

Not all of these carry equal weight. A rough severity guide:

SignSeverityFirst corrective step
Shifting success measureHighRe-agree one measurable goal in writing with the budget holder
No live or recent-data resultsHighTest on last month's data this week
No end-user contactHighPut ten real predictions in front of users now
Stalled data accessMedium to highEscalate with a named owner and a date
Label disagreementMediumRun a labelling session and write the rules down
Too-good resultsMediumAudit features for leakage before celebrating
Tiny accuracy gainsLow to mediumFreeze the model and work on deployment
No evaluation setMediumBuild one from 200 real examples

Three or more high-severity signs together usually mean the project needs a pause and a proper review rather than another sprint.

One caution about using this list. A single sign on its own is normal. Every machine learning project has a week where the data access stalls or a result looks suspiciously good. The pattern to worry about is persistence: the same sign appearing in three consecutive updates, or several signs arriving together. Keep a simple note of which ones you see at each review and the trend becomes obvious long before anyone needs to call it a failure.

What to do this week

If you recognise your project, three moves are cheap and informative. First, ask for performance on the most recent month of data, not the test set. Second, put a sample of real predictions in front of the people who would use them and watch their reaction. Third, ask the team to write one sentence describing what 'done' means.

The answers usually tell you whether you have a project that needs refocusing or one that needs stopping. Both are better outcomes than drifting. If it is the former, our guide to rescuing a stalled machine learning project sets out the audit we run.

An outside view helps

Teams inside a struggling project are usually the last to see it, not through dishonesty but because they are close to the daily progress and far from the original business case. SpiderHunts is regularly asked for a short independent review of machine learning and AI projects run by other suppliers or in-house teams. The most common finding is not bad engineering. It is a project that lost track of the decision it was meant to improve.

Frequently asked questions

How do you know if an AI project is failing?

Watch for a changing definition of success, results only reported on test data, no involvement from end users and data problems that never quite get resolved. Any two of those together deserve a serious conversation.

What percentage of machine learning projects fail?

Figures quoted vary widely and are rarely measured consistently, so we would not trust a single number. What is clear is that many projects never reach production, and most of the causes are organisational rather than technical.

Should we stop an AI project that is behind schedule?

Not automatically. Check whether there is evidence of value on recent data and a credible path to deployment. If both exist, refocus; if neither does after several months, stopping is often the cheaper decision.

Why do AI results look great in testing and poor in production?

The most common causes are data leakage, testing on a random sample instead of recent data, and production data that is messier or arrives later than the training data assumed.

Keep reading

Seeing a few of these signs on your project?

Tell us what is happening. We will give you a straightforward second opinion on how serious it is and what we would change first, whoever is doing the work.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

Custom Software DevelopmentDigital Transformation