Designing AI Projects Around Business Outcomes
Last updated:
Start from the number that should move
Most AI projects are described by what will be built: a chatbot, a document reader, a forecasting model. We ask a different first question. When this is working, which number in the business will be different, and who looks at that number today?
If nobody can answer, the project is at risk before it starts. It will be judged on impressions, and impressions fade the first time the system makes a visible mistake. If someone can answer, everything else follows: what to build, how to test it, and whether it was worth the money.
Outcomes that suit AI well
| Outcome | Typical AI approach | How it is measured |
|---|---|---|
| Less time per task | Extraction, classification, drafting | Minutes per case, sampled before and after |
| Fewer errors | Validation plus flagged review | Errors found downstream per hundred cases |
| Faster response | Routing and first-draft replies | Median time to first meaningful response |
| Fewer lost enquiries | Out-of-hours handling and triage | Enquiries unanswered or abandoned |
| Faster cash collection | Invoice processing, payment risk scores | Days from delivery to invoice or payment |
| Fewer missed risks | Anomaly detection, document review | Issues caught before rather than after impact |
Each of these is something a finance director or operations manager already understands. That matters, because they will decide whether the next project gets funded.
Outputs are not outcomes
Model accuracy, number of questions answered and documents processed are outputs. They matter, but they are not why a business pays. A chatbot can answer thousands of questions and move no business number at all if the questions were ones customers could already answer from the website.
We track outputs as leading indicators, because outcomes can take weeks to show up. Acceptance rates, coverage of eligible cases and escalation rates tell you early whether the system is on track. The outcome tells you whether it was worth it.
Nobody has ever renewed a budget because a model reached a particular F1 score. They renew it because the backlog went away.
Writing the outcome into the scope
When SpiderHunts scopes an AI project, the outcome is written into the document alongside the features. It has five parts, and we agree each one with you.
- The outcome: in plain words, such as 'time spent keying supplier invoices'
- The baseline: measured, not estimated, over a normal period
- A realistic range: what a good result would look like, stated as a range
- The measurement method: who measures, how, and from which data
- The check date: when the outcome will be reviewed after launch
An illustrative example: a 30-person accountancy practice spends roughly 60 staff hours a week extracting figures from client documents. The target range might be to reduce that by a third to a half within three months of launch, measured by timesheet codes the practice already uses, reviewed at the end of the quarter. Those numbers are an example, not a promise; the point is that they are written down before anyone builds.
Outcomes we will not promise
Some outcomes are sensible to hope for and wrong to guarantee. We do not promise headcount reductions, because in our projects the typical result is a change in what people do, and a business case built on redundancies tends to meet resistance that sinks the project. We do not guarantee revenue increases, because too many factors outside the system affect them.
Outcome-based pricing, where the supplier is paid by results, is getting more attention in 2026 and can work where the outcome is narrow, measurable and mostly within the system's control. For most bespoke business projects it is not, and it tends to produce arguments about attribution. We explain our preference for fixed-price scoping in fixed price versus time and materials.
When the outcome does not move
Sometimes the system works as designed and the number does not change. That is uncomfortable and extremely informative. The usual causes are that the time saved was absorbed elsewhere, the task was not really the bottleneck, or staff are not using the feature for most cases.
- Check usage first: is the system handling the eligible cases?
- Check the bottleneck: did work simply pile up at the next step?
- Check the measurement: is the baseline comparable to the new period?
- Decide honestly whether to adjust, extend or retire the feature
Our enterprise AI engagements include this review as a scheduled step, so a disappointing result gets investigated rather than quietly ignored.
Why this protects you more than us
Designing around outcomes makes our job harder in one way: it gives you a clear standard to hold us to. We think that is right. It also makes the next conversation with your board, your investors or your finance team far easier, because you can show a number that moved rather than a demo that impressed.
It changes the build too. When the outcome is time saved on invoice keying, we spend effort on the awkward suppliers that cause most of the delay, not on a polished interface for the easy ones. When the outcome is fewer lost enquiries, the out-of-hours path gets the attention, not the daytime chat widget. The outcome decides where the engineering hours go, which is exactly where a buyer wants them.
Frequently asked questions
How do you measure the ROI of an AI project?
What if we have no baseline data?
Does SpiderHunts offer outcome-based pricing for AI projects?
How long after launch should we expect to see business outcomes?
Which business outcome should our first AI project target?
Know which number you want AI to move?
Tell us the figure and how it is measured today. We will tell you whether AI can plausibly move it, by roughly how much in an illustrative case, and what we would measure first.
Related services
What we build for problems like this one