Signs Your AI Project Is Heading for Trouble
Last updated:
Projects do not fail suddenly
AI projects rarely end with a dramatic failure. They drift. Updates get vaguer, the go-live date moves by a month, then another, and eventually the budget holder stops asking. By then most of the money is gone.
The signs are there much earlier. Some are general project warnings, which we covered in signs your software project is in trouble. The ones below are specific to machine learning and AI work, where the uncertainty is higher and the problems hide more easily behind technical language.
Warning signs in how progress is reported
- The success measure has changed more than once. It started as 'reduce manual review by half', became 'improve accuracy', and is now 'build a foundation for future AI'. Each change moves the goalposts closer to wherever the team happens to be.
- Every number is a test-set score. Months in, nobody can show performance on live or recent data. Test scores are easy to make look good.
- Updates describe activity, not results. 'We tried three new architectures' is activity. 'Error on last month's invoices fell from 12% to 8%' is a result.
- Accuracy keeps improving by tiny amounts. Weeks spent moving from 88.1% to 88.6% usually means the team is tuning because the real blockers, such as data or integration, are harder to face.
Warning signs in the data
- Data issues are always 'nearly sorted'. If access to a key source has been a week away for two months, something political or technical is wrong and needs escalating.
- Labels are being argued about. Two experts disagree about what the right answer is for a quarter of examples. The model cannot learn a distinction the business has not agreed on.
- Results look too good. A sudden jump to 98% is more often data leakage than a breakthrough. Treat surprise good news with the same suspicion as bad.
- The training data is old or hand-picked and nobody has tested on what arrived last month.
Warning signs in people and process
- No end user has touched it. If the people who will rely on the model have not seen a real prediction by the halfway point, adoption will be the next crisis.
- The decision owner has gone quiet, stopped attending reviews or delegated them to someone junior. Interest has moved on, and so will funding.
- Nobody has asked how it will be deployed. Integration is still described as 'the easy part at the end'. It rarely is.
- For language model projects, there is no evaluation set. Quality is judged by someone reading a few outputs and saying they look good. Every prompt change is a gamble.
How serious is it?
Not all of these carry equal weight. A rough severity guide:
| Sign | Severity | First corrective step |
|---|---|---|
| Shifting success measure | High | Re-agree one measurable goal in writing with the budget holder |
| No live or recent-data results | High | Test on last month's data this week |
| No end-user contact | High | Put ten real predictions in front of users now |
| Stalled data access | Medium to high | Escalate with a named owner and a date |
| Label disagreement | Medium | Run a labelling session and write the rules down |
| Too-good results | Medium | Audit features for leakage before celebrating |
| Tiny accuracy gains | Low to medium | Freeze the model and work on deployment |
| No evaluation set | Medium | Build one from 200 real examples |
Three or more high-severity signs together usually mean the project needs a pause and a proper review rather than another sprint.
One caution about using this list. A single sign on its own is normal. Every machine learning project has a week where the data access stalls or a result looks suspiciously good. The pattern to worry about is persistence: the same sign appearing in three consecutive updates, or several signs arriving together. Keep a simple note of which ones you see at each review and the trend becomes obvious long before anyone needs to call it a failure.
What to do this week
If you recognise your project, three moves are cheap and informative. First, ask for performance on the most recent month of data, not the test set. Second, put a sample of real predictions in front of the people who would use them and watch their reaction. Third, ask the team to write one sentence describing what 'done' means.
The answers usually tell you whether you have a project that needs refocusing or one that needs stopping. Both are better outcomes than drifting. If it is the former, our guide to rescuing a stalled machine learning project sets out the audit we run.
An outside view helps
Teams inside a struggling project are usually the last to see it, not through dishonesty but because they are close to the daily progress and far from the original business case. SpiderHunts is regularly asked for a short independent review of machine learning and AI projects run by other suppliers or in-house teams. The most common finding is not bad engineering. It is a project that lost track of the decision it was meant to improve.
Frequently asked questions
How do you know if an AI project is failing?
What percentage of machine learning projects fail?
Should we stop an AI project that is behind schedule?
Why do AI results look great in testing and poor in production?
Seeing a few of these signs on your project?
Tell us what is happening. We will give you a straightforward second opinion on how serious it is and what we would change first, whoever is doing the work.
Related services
What we build for problems like this one