Measuring the Productivity of AI Agents
Last updated:
The dashboard that says everything is wonderful
Six months after launch, the agent dashboard shows 42,000 tasks handled, 1.3 million tokens processed and a satisfaction score of 4.6. The board is pleased. Then the operations director points out that headcount in the team is unchanged, overtime is up, and the backlog is where it was.
Both can be true. The agent may be handling tasks that were not the bottleneck, producing drafts that take nearly as long to fix as to write, or creating new work downstream. Measuring AI agent productivity properly means measuring what happened to the work, not what the agent did.
Activity metrics vs outcome metrics
| Tells you it is busy | Tells you it is useful |
|---|---|
| Tasks handled or touched | Tasks completed without human rework |
| Messages or tokens processed | Cost per completed task, including review time |
| Automation rate as claimed by the tool | Share of cases fully resolved, checked by sampling |
| Average response time of the agent | End-to-end time from request to resolution |
| Satisfaction rating on the chat widget | Repeat contacts, complaints and reversals |
Activity metrics are still worth collecting for operations and cost control. They just should not appear on a slide claiming the agent is productive. Our post on agent observability covers the operational side.
The five numbers we would put in front of a board
- Cost per completed task. Model and infrastructure costs plus the staff time spent reviewing, correcting and handling escalations, divided by tasks genuinely finished.
- Clean completion rate. The share of agent outputs accepted without edits or rework, measured by sampling or by tracking edits.
- Cycle time. Time from a request arriving to it being resolved, compared with the same measure before the agent.
- Error and reversal rate. Refunds reissued, orders corrected, tickets reopened, journals reversed that trace back to agent actions.
- Capacity released. Hours freed, and crucially what those hours went on instead, backed by something observable such as backlog shrinking.
Number five is the one most often fudged. Saved hours that disappear into general busyness are not savings. If the capacity was meant for something specific, measure that thing.
Setting a baseline before the agent launches
Without a baseline, every productivity claim is a guess. The baseline does not need to be elaborate, but it must exist before launch.
- Record volume, cycle time and error rate for the process for at least four weeks
- Time a sample of tasks done by people, including the awkward ones, rather than relying on estimates
- Note the cost of the people and tools currently involved
- Capture quality measures you already have, such as reopen rates or audit findings
- Write down what the agent is expected to change, in numbers, so the review has something to test
A useful illustration: a claims team processes 900 claims a month with an average of 35 minutes of handling each. After an agent prepares each claim, handling falls to 20 minutes, but 8% of prepared claims need a rework that takes 25 minutes. The saving is real, but smaller than the headline figure of 15 minutes, and that is the number that belongs in the business case. The same arithmetic appears in our guide to calculating automation ROI.
Hidden costs that erase agent productivity
- Review fatigue. Checking agent output is tiring, and reviewers either slow down or stop checking properly.
- Downstream rework. A wrong classification costs little to the agent and a lot to the team that receives it.
- Exception concentration. The agent takes the easy cases, so people handle only hard ones, and per-task time for people rises.
- Maintenance. Prompt updates, evaluation runs, integration fixes and vendor changes all take time.
- Supervision. Someone has to own the agent, read its logs and answer for its mistakes.
Exception concentration catches many teams out. If average handling time for people rises after launch, that is often expected rather than a failure, and the productivity measure should be for the whole process rather than per person.
Measuring fairly: comparison groups and sampling
Where you can, compare like with like. Running the agent on a random half of incoming work for a few weeks, with people handling the other half as before, gives a far clearer answer than comparing this quarter with last. Seasonality, staffing changes and product launches all distort before-and-after comparisons.
Quality needs sampling. Pick a fixed number of agent-completed tasks every week, have someone who knows the work mark them as correct, acceptable with edits or wrong, and track the trend. It takes an hour a week and is the most honest productivity signal available. Treat it as ongoing evaluation rather than a launch activity.
How we report on agents we build
SpiderHunts agrees the baseline and the five numbers with the client before building. The agent logs what it needs to calculate them from the start, and the monthly report shows cost per completed task and clean completion rate first, activity figures last. When a number goes the wrong way, the report says so.
If you want help setting up that measurement for an agent you already run, our data science team does this kind of analysis. For the wider evidence on AI and team output, see whether AI makes teams more productive.
Frequently asked questions
How do you measure AI agent productivity?
What is a good automation rate for an AI agent?
How long before we can judge an agent's productivity?
Should human review time count against the agent?
Can we trust the productivity figures in a vendor's dashboard?
Not sure whether your agent is paying off?
Send us what you currently track about your agent. We will tell you which numbers mean something and what to add so the next review has a real answer.
Related services
What we build for problems like this one