How to measure the ROI of AI agents
Measure AI agent ROI per piece of work: take a baseline for quality, cost, time and human effort, track cost per accepted task and agree when to stop or scale. Formulas and a worked example.
To measure the ROI of an AI agent, measure a piece of work, not the agent. Record a baseline for quality, cost, time and human effort before anything is built. After launch, track the cost per accepted task and compare it with the baseline. The return is what that change is worth over a period, minus what it cost to get there, divided by that cost.
Most companies skip the baseline, and it shows. In McKinsey’s 2025 survey, 88% of companies use AI in at least one function and 39% report any EBIT impact. MIT’s 2025 study found 95% of organisations see no measurable P&L impact from their generative AI pilots. Part of that gap is measurement: if nobody wrote down how the work performed before, nobody can say what changed.
1. Pick the unit of work
ROI needs something to count. Pick the smallest unit of work that has a clear result: one matched invoice, one meeting brief, one reconciled account, one enriched contact. Then define what accepted means for that unit. An accepted task meets the agreed standard without being redone. A cheap run that a person has to redo isn’t cheap.
2. Take a baseline before you build
Measure the work as it runs today on the same four things the agent will be measured on. We agree these on real examples during the audit, as described in agree the standard before you build.
| Measure | The question | How to record the baseline |
|---|---|---|
| Quality | Does the work meet the standard on real examples? | Check a sample of today’s output against the agreed answers |
| Cost | What does each completed piece of work cost? | Hours per task times the loaded hourly cost, plus tools and outside data |
| Time | How long before the result is ready? | Elapsed time from the trigger to a usable result |
| Human effort | What review or correction is still needed? | Minutes of review per task, and the share of tasks sent back |
3. Count the cost per accepted task
After launch, add up everything it takes to get the work done to the standard, including the people who still review it, and divide by the tasks that passed.
- Cost per accepted task = (agent run costs + platform and service costs for the period + review hours times hourly cost) ÷ accepted tasks
- Acceptance rate = accepted tasks ÷ tasks attempted
Leaving out review time is the most common way to make an agent look better than it is. Leaving out the cost of the work that now gets done properly, and wasn’t before, is the most common way to make it look worse.
4. Put a value on the change
Value comes from three places. Count each one only when someone owns the number.
- Time released. Only count hours that go to other work. If the time isn’t redeployed, it is a capacity gain, not a saving.
- Faster results. Count it when something depends on the speed, such as an invoice paid on time or a proposal out before a competitor’s.
- Better quality. Fewer errors, fewer missed exceptions, wider coverage. Put a number on it only where the cost of an error is known.
Then use three formulas:
- Net monthly value = (baseline cost per task minus cost per accepted task) × accepted tasks per month
- Payback period in months = one-off costs ÷ net monthly value
- ROI over a period = (net value over the period minus one-off costs) ÷ one-off costs
Running costs are already inside the cost per accepted task, so the one-off costs here are the audit, the build and any setup on your side.
5. A worked example, with made-up numbers
Every number in this example is hypothetical. It shows the arithmetic and does not come from a client. The platform and service figure is a placeholder, not our pricing.
Say a finance team checks 2,000 invoices a month against orders. Today each check takes 6 minutes at a loaded cost of €60 an hour, so the baseline is €6.00 per invoice, or €12,000 a month. After the build, an agent checks every invoice. 85% match and pass on their own. The other 300 go to a person with the evidence attached, and each takes 4 minutes to resolve.
| Line | Hypothetical figure | Working |
|---|---|---|
| Agent run costs | €200 a month | €0.10 × 2,000 invoices |
| Platform and service | €4,000 a month | Placeholder |
| Review time | €1,200 a month | 300 exceptions × 4 minutes = 20 hours × €60 |
| Total running cost | €5,400 a month | €200 + €4,000 + €1,200 |
| Cost per accepted task | €2.70 | €5,400 ÷ 2,000 |
| Net monthly value | €6,600 | (€6.00 minus €2.70) × 2,000 |
| One-off costs | €30,000 | Placeholder for audit and build |
| Payback period | About 4.5 months | €30,000 ÷ €6,600 |
| ROI over year one | 164% | (12 × €6,600 minus €30,000) ÷ €30,000 |
This only counts time, and assumes the released hours go to other work. It leaves out any value from fewer payment errors or invoices paid on time, which in a real case could matter more. A real, anonymised baseline from an engagement would make this example stronger, and we will add one when we can share it.
6. Agree when to stop, fix or scale
Agree the thresholds before the build, at the same time as the standard. Then the decision after launch is a reading, not a debate.
| What you see | What to do |
|---|---|
| Quality meets the standard and cost per accepted task is below the baseline for several cycles | Scale it: more volume, or the next piece of work |
| Quality meets the standard but cost is too high | Fix it: narrow the context, change the model or move steps back to ordinary software |
| Quality is close but review time is high | Fix it: improve the evaluations and the cases that go to people |
| Quality is still short after the agreed number of cycles | Stop it, or redesign the work before trying again |
Mistakes that skew AI agent ROI
- Measuring the agent’s output instead of the work. A good answer that nobody uses has no return.
- Counting the model bill and forgetting review time, platform costs and the build.
- Counting hours saved that nobody redeploys.
- Having no baseline, so every result is compared with a guess.
- Measuring once at launch. The company keeps changing, and the numbers drift with it unless the evaluations keep running.
How we measure it at Ortelian
We agree the baseline in the two-week audit on site, before anything is built, and the first job goes live when it passes the evaluations we agreed. Detail on how we work. The platform records every run and the version it used, so the counts behind cost per accepted task come from records rather than estimates. After launch, your team runs the work or we keep running it for you. Either way, we host and operate the platform and the evaluations keep running.
Measurement sits next to the other controls in AI agent governance. To check whether a piece of work is ready to measure at all, start with the AI readiness assessment. For the full sequence, see how to deploy AI agents in an enterprise.
Questions people ask
How do you calculate the ROI of an AI agent?
Take a baseline cost per task before the build. After launch, divide everything it takes to get the work done to the standard, including review time, by the tasks that passed. Multiply the difference by monthly volume to get net monthly value. ROI over a period is that value minus the one-off costs, divided by the one-off costs.
What is a good ROI for AI agents?
There is no honest general number. It depends on the volume, the cost of the work today and what an error costs. Agree the threshold for your own work before you build, and judge the agent against it.
How long does it take to see a return?
Measure from the day the first job passes its evaluations and goes live. The payback period then follows from the one-off costs and the net monthly value. In the hypothetical example above it was about four and a half months, but your numbers will differ.
Should we count hours saved?
Only the hours that go to other work. Otherwise count them as capacity, and say so.
