Your First 30 Days on an AI Project
Last updated:
Why the first month of an AI project is different
We have described the first month of an ordinary software engagement in what happens when you hire SpiderHunts. An AI project follows the same principles of fixed scope and visible progress, but the first month has a different priority.
In a normal build, the main uncertainty is effort. In an AI build, the main uncertainty is whether the approach works well enough on your data at all. So the first 30 days are arranged to answer that question as cheaply as possible, before anyone spends serious money on screens, integrations and polish.
Days 1 to 5: real examples and real access
The most valuable thing you can give us in the first week is a pile of real, messy examples of the task. Not the tidy ones chosen for a presentation. The scanned invoice at an angle, the customer email written in three languages, the order with handwritten amendments.
- Several hundred real examples of the input, anonymised where necessary
- The correct output for as many as you have, such as the order as it was eventually keyed
- Read access to the systems the AI will eventually read from or write to
- An hour with the person who does the task, sharing their screen
- Your data protection lead's requirements, so the design fits them from the start
Access is the item most likely to slip. Start arranging it the day the project is agreed.
Days 6 to 12: the evaluation set and the baseline
Before we build anything clever, we build the test. With the person who knows the task, we assemble an evaluation set: a few hundred real cases with agreed correct answers, deliberately including the difficult ones. Every version of the system from now on is scored against it.
At the same time we measure the current process. How long does a case take? How many errors reach the next step? That baseline is what the AI has to beat, and it is written down now so nobody can move the goalposts later, including us.
The evaluation set is the most important thing built in the first month. The prototype is replaceable. The test is not.
Days 13 to 25: a working slice on real data
Now we build. The goal is a thin, working version of the core task running against your real examples. For a document extraction project, that means documents in, structured fields out, with a confidence score per field. For a support triage feature, tickets in, categories and draft replies out. Our AI integration service describes this as a proof of concept on real data, typically within two to three weeks.
It is deliberately unpolished. There may be no proper interface, just results in a spreadsheet or a simple review screen. What matters is the score on the evaluation set, the cost per case at realistic volume, and the list of cases it gets wrong. We share those results as they improve, not only at the end.
Days 26 to 30: the go or no-go review
The month ends with a review meeting and a short written report. It sets out what the system achieved on the evaluation set, how that compares with the baseline, the estimated running cost, which kinds of cases it struggles with, and our recommendation.
| What the evidence shows | What it usually means | Recommended next step |
|---|---|---|
| Strong results on most cases, clear failure patterns | The approach works | Proceed to the production build with review routing for hard cases |
| Good results on some case types only | Works for part of the task | Narrow the scope to the cases that work |
| Results depend on data you do not capture | The data is the constraint | Fix data capture first, revisit later |
| Results close to or below the baseline | AI is not the right tool here | Stop, or solve with conventional software |
We would much rather recommend stopping at day 30 than at month four. So would you.
What we need from you each week
- Week one: examples, access and an hour with the person who does the task
- Week two: a few hours from that person to agree correct answers for difficult cases
- Week three: quick answers when we find cases where even your team disagree on the right output
- Week four: the decision-maker at the review meeting
The third item surprises people. Building an evaluation set often reveals that two experienced staff handle the same case differently. Settling that is valuable in its own right, and the AI cannot be more consistent than the rule it is given.
Common surprises in month one
The data is messier than anyone described, usually in ways that matter. A small number of case types account for most of the difficulty. The running cost is either much lower than feared or needs a routing design to keep it sensible. And the person who does the task often has ideas for improvement that are worth more than the original brief. If you want a view of the whole engagement beyond this first month, what SpiderHunts brings to machine learning projects covers what you own at the end.
Frequently asked questions
How long before we see a working AI prototype?
How many examples do we need to start an AI project?
What happens if the first month shows AI will not work?
Do we need to prepare a detailed specification before the first month?
Is our data safe during the first month?
Ready to find out if your AI idea survives real data?
Tell us the task and send a handful of real examples. We will tell you what the first month would test and what you would need to have ready.