Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI Apps

What Actually Happens During an Eight-Week AI Project

Last updated:

Weeks 1–2: defining correct

We build the evaluation set with your domain expert: a hundred or more real cases with agreed right answers. This is the least exciting fortnight and the one that decides the project.

In parallel we establish a baseline — what accuracy a straightforward approach achieves — so every later improvement is measured against something.

Weeks 3–4: pipeline and integration

  • The extraction or generation pipeline against real data
  • Validation rules from your master data and business logic
  • Integration with the destination system, including idempotency
  • Confidence scoring and the routing thresholds

By the end of week 4 there is something producing output, measured against the evaluation set, with a number attached.

Weeks 5–6: the interface and the guardrails

This is where a demo becomes a system. The review interface, the failure behaviours, the cost caps, the monitoring, the audit trail and the switch that turns it off.

We demonstrate the review interface to the people who will actually use it in week 5, and change it in week 6 based on what they say.

Weeks 7–8: shadow and launch

  1. Shadow running: the system processes real work nobody acts on
  2. Comparison against what humans actually did, case by case
  3. Threshold tuning based on that comparison
  4. Launch to two or three users, with the switch in reach
  5. Daily review for the first week, then weekly

Your time, concentrated

WeekFrom youRoughly
1–2Domain expert on the evaluation set2–3 days
3–4Questions answered, access provided2 hours
5–6Interface review with real usersHalf a day
7Shadow results reviewed togetherHalf a day
8Launch decision1 hour

Booked in advance, that is manageable. Assumed as ad hoc availability, it is the reason projects slip.

Frequently asked questions

What if week 2 shows it will not work?

We tell you then, and you have spent a fortnight rather than a quarter. That is the point of measuring first.

Can it go faster?

Sometimes, by narrowing scope rather than compressing steps. Removing the evaluation set to save a fortnight removes the ability to know whether it works.

What happens after week 8?

Four to six weeks of tuning as real usage reveals real edge cases, then a support arrangement. Quality settles about six weeks after first users, not on launch day.

Do you work on site?

Mostly remotely, with on-site sessions where they help — usually the evaluation set and the interface review.

Keep reading

Want to know what your project would look like week by week?

Tell us the process and we will map it out, with the dates we would need your people.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsCustom Software DevelopmentSaaS Development