Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI Apps

Measuring the Productivity of AI Agents

Last updated:

The dashboard that says everything is wonderful

Six months after launch, the agent dashboard shows 42,000 tasks handled, 1.3 million tokens processed and a satisfaction score of 4.6. The board is pleased. Then the operations director points out that headcount in the team is unchanged, overtime is up, and the backlog is where it was.

Both can be true. The agent may be handling tasks that were not the bottleneck, producing drafts that take nearly as long to fix as to write, or creating new work downstream. Measuring AI agent productivity properly means measuring what happened to the work, not what the agent did.

Activity metrics vs outcome metrics

Tells you it is busyTells you it is useful
Tasks handled or touchedTasks completed without human rework
Messages or tokens processedCost per completed task, including review time
Automation rate as claimed by the toolShare of cases fully resolved, checked by sampling
Average response time of the agentEnd-to-end time from request to resolution
Satisfaction rating on the chat widgetRepeat contacts, complaints and reversals

Activity metrics are still worth collecting for operations and cost control. They just should not appear on a slide claiming the agent is productive. Our post on agent observability covers the operational side.

The five numbers we would put in front of a board

  1. Cost per completed task. Model and infrastructure costs plus the staff time spent reviewing, correcting and handling escalations, divided by tasks genuinely finished.
  2. Clean completion rate. The share of agent outputs accepted without edits or rework, measured by sampling or by tracking edits.
  3. Cycle time. Time from a request arriving to it being resolved, compared with the same measure before the agent.
  4. Error and reversal rate. Refunds reissued, orders corrected, tickets reopened, journals reversed that trace back to agent actions.
  5. Capacity released. Hours freed, and crucially what those hours went on instead, backed by something observable such as backlog shrinking.

Number five is the one most often fudged. Saved hours that disappear into general busyness are not savings. If the capacity was meant for something specific, measure that thing.

Setting a baseline before the agent launches

Without a baseline, every productivity claim is a guess. The baseline does not need to be elaborate, but it must exist before launch.

  • Record volume, cycle time and error rate for the process for at least four weeks
  • Time a sample of tasks done by people, including the awkward ones, rather than relying on estimates
  • Note the cost of the people and tools currently involved
  • Capture quality measures you already have, such as reopen rates or audit findings
  • Write down what the agent is expected to change, in numbers, so the review has something to test

A useful illustration: a claims team processes 900 claims a month with an average of 35 minutes of handling each. After an agent prepares each claim, handling falls to 20 minutes, but 8% of prepared claims need a rework that takes 25 minutes. The saving is real, but smaller than the headline figure of 15 minutes, and that is the number that belongs in the business case. The same arithmetic appears in our guide to calculating automation ROI.

Hidden costs that erase agent productivity

  • Review fatigue. Checking agent output is tiring, and reviewers either slow down or stop checking properly.
  • Downstream rework. A wrong classification costs little to the agent and a lot to the team that receives it.
  • Exception concentration. The agent takes the easy cases, so people handle only hard ones, and per-task time for people rises.
  • Maintenance. Prompt updates, evaluation runs, integration fixes and vendor changes all take time.
  • Supervision. Someone has to own the agent, read its logs and answer for its mistakes.

Exception concentration catches many teams out. If average handling time for people rises after launch, that is often expected rather than a failure, and the productivity measure should be for the whole process rather than per person.

Measuring fairly: comparison groups and sampling

Where you can, compare like with like. Running the agent on a random half of incoming work for a few weeks, with people handling the other half as before, gives a far clearer answer than comparing this quarter with last. Seasonality, staffing changes and product launches all distort before-and-after comparisons.

Quality needs sampling. Pick a fixed number of agent-completed tasks every week, have someone who knows the work mark them as correct, acceptable with edits or wrong, and track the trend. It takes an hour a week and is the most honest productivity signal available. Treat it as ongoing evaluation rather than a launch activity.

How we report on agents we build

SpiderHunts agrees the baseline and the five numbers with the client before building. The agent logs what it needs to calculate them from the start, and the monthly report shows cost per completed task and clean completion rate first, activity figures last. When a number goes the wrong way, the report says so.

If you want help setting up that measurement for an agent you already run, our data science team does this kind of analysis. For the wider evidence on AI and team output, see whether AI makes teams more productive.

Frequently asked questions

How do you measure AI agent productivity?

Measure outcomes against a baseline: cost per completed task including human review, share of work completed without rework, end-to-end cycle time, and errors or reversals caused by the agent. Activity counts alone are not evidence of productivity.

What is a good automation rate for an AI agent?

There is no universal figure. A lower automation rate with very few errors can be worth more than a high one that creates rework. Judge it by cost per completed task and quality on sampled cases.

How long before we can judge an agent's productivity?

Usually two to three months after launch, once early fixes are done and people have settled into the new process. Collect a baseline for at least four weeks before launch.

Should human review time count against the agent?

Yes. Review, correction and escalation time are part of what the agent costs to run. Leaving them out is the most common way agent productivity gets overstated.

Can we trust the productivity figures in a vendor's dashboard?

Treat them as activity figures unless the vendor shows how outcomes are verified. Check them against your own sampled quality reviews and business measures such as backlog and cycle time.

Keep reading

Not sure whether your agent is paying off?

Send us what you currently track about your agent. We will tell you which numbers mean something and what to add so the next review has a real answer.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsCustom Software DevelopmentSaaS Development