Choosing Your First AI Project by Risk and Reward
Last updated:
The trouble with the idea list
Most businesses that decide to 'do something with AI' end up with a list of ideas within a fortnight. A customer chatbot. Automated quoting. Forecasting. Summarising calls. Something with the contracts. The list is not the hard part.
The hard part is that people pick the idea with the biggest potential, which is almost always also the riskiest. A failed first project does more damage than a modest successful one does good, because it sets the organisation's view of AI for the next two years.
Four scores for every idea
Score each idea from 1 to 5 on four dimensions. It takes an hour with the right people in the room.
- Value. Hours saved, revenue protected or errors avoided per year. Rough numbers are fine. 5 is significant; 1 is nice to have.
- Cost of a wrong answer. Scored in reverse: 5 means a mistake is cheap and caught by a person, 1 means it reaches a customer, a regulator or the accounts unchecked.
- Data readiness. 5 means the examples exist, are accessible and are roughly clean. 1 means they need collecting, labelling or extracting from paper.
- Effort. Scored in reverse: 5 means one system and a short build, 1 means several integrations, new processes and change management across teams.
Multiply value by the average of the other three, or simply add them. The exact formula matters less than forcing the conversation about each dimension.
A worked example
Here is how an illustrative 60-person wholesale distributor's list might score:
| Idea | Value | Error cost (rev.) | Data | Effort (rev.) | Total |
|---|---|---|---|---|---|
| Customer-facing chatbot for order queries | 3 | 2 | 3 | 2 | 10 |
| Extract purchase orders from emailed PDFs | 4 | 4 | 5 | 4 | 17 |
| Automated pricing recommendations | 5 | 1 | 3 | 2 | 11 |
| Demand forecast for top 200 products | 4 | 3 | 4 | 3 | 14 |
| Summarise sales calls into the CRM | 2 | 5 | 4 | 4 | 15 |
Purchase order extraction wins clearly. It is dull. It is also frequent, measurable, tolerant of a review step and built on documents the business already receives. Pricing has the highest value and the worst risk profile, which is exactly why it should be the third or fourth project, once the team knows how these systems behave.
Why cost of error deserves extra weight
If you only change one habit, weight the error score heavily. AI systems are wrong some of the time; the only question is what happens when they are. A project where a person naturally reviews the output, such as a draft, a suggestion or a pre-filled form, gets real value from 85% accuracy. A project where the output acts directly needs far more engineering to be safe, and the first project is the wrong place to learn how to do that.
The best first AI project is one where being wrong is boring.
Ideas that look good and score badly
- The public chatbot. Highly visible, so a mistake is embarrassing, and often less valuable than assumed if most queries already have a self-service answer.
- Anything needing new data collection. If labels do not exist yet, the project is really a data project with a model at the end.
- Cross-department process change. Technically simple, organisationally slow. Leave it until you have a success to point to.
- The CEO's pet idea with no clear owner below the CEO. Nobody has time to make it work day to day.
Reward is more than hours saved
Some projects score modestly on direct value but pay off in other ways. A first project that forces a clean data pipeline out of the ERP makes the second and third projects much cheaper. One that gets a sceptical operations team using and trusting a model changes what they will agree to next.
Add a note beside each idea about what it would make easier afterwards. It rarely changes the top pick, but it often breaks ties. There are more ideas about good starting points in our post on first AI project recommendations.
What to do with the ideas that did not win
Do not throw the list away. The ideas that scored badly usually scored badly for a specific, fixable reason, and writing that reason down turns a rejected idea into a plan.
- If data readiness was the weak score, the fix may be a change to how information is captured today, so that in a year the examples exist
- If cost of error was the weak score, ask whether a review step could be added, turning an automatic decision into a suggestion
- If effort was the weak score, check whether the first project's pipeline or integration would reduce it
- If value was the weak score, it may simply not be worth doing, and that is a useful conclusion too
Revisit the scores every six months. A distributor that has shipped purchase order extraction now has clean order data flowing automatically, which quietly lifts the data score of the demand forecast. Projects compound in a way the first scoring session cannot see.
Running the session
Get the people who own the processes, not only the people interested in AI. The warehouse manager knows whether a wrong forecast costs a missed delivery or a slightly overfull shelf, and that answer changes the score.
SpiderHunts runs this as a half-day workshop, usually ending with one project to scope properly and one or two to revisit in six months. If the top idea involves agents or multi-step automation, our AI agents work uses the same risk thinking, with approval steps designed around the actions that score worst on cost of error.
Frequently asked questions
What makes a good first AI project?
Should our first AI project be customer-facing?
How do you estimate the value of an AI project?
How many AI projects should we start at once?
Have a list of AI ideas and no clear first pick?
Send us the list. We will score it with you using the method below and tell you which idea we would start with, and which we would leave alone for now.
Related services
What we build for problems like this one