# How do we measure ROI on an AI project?

_Buying AI services · Last updated 2026-08-20 · Bitfumes_

## Short answer

Name one business metric per use case, record its baseline before any build starts, and write the payback arithmetic down while it can still change the decision. Measure against a holdout — a comparable team, queue or period still working the old way — because a before-and-after comparison absorbs every other change your business made that quarter. Count the full cost, which means build, inference at real volume, and the ongoing maintenance an AI system needs to resist drift, not just the invoice. In back-office automation a use case that cannot show payback within twelve months usually has a scope problem rather than a technology problem.

## Four numbers, in this order

- The baseline — what the metric reads today, taken from a system you already trust, before anyone starts building. This is the number that is almost always missing, and it cannot be reconstructed afterwards.
- The target — what the metric needs to read for this to have been worth doing, agreed with whoever owns the budget.
- The total cost — build, inference at projected volume, human review time, and maintenance for the first year.
- The measurement method — who reads the metric, how often, and against what comparison group. Decided before launch, because deciding afterwards is how a disappointing result becomes a debate about methodology.

## Pick the metric that cannot be gamed

| Use case | Metric that works | Metric that misleads |
| --- | --- | --- |
| Support triage and drafting | Median first-response time; share resolved without escalation | Tickets 'touched by AI' |
| Document intake and extraction | Cost per document processed; error rate against manual handling | Documents processed |
| Quote and proposal generation | Quote turnaround time; win rate on quotes returned within a day | Quotes generated |
| Internal knowledge search | Time to find an answer; volume of repeat questions to experts | Searches run |
| Code assistance | Cycle time to merge; change failure rate | Lines of suggested code accepted |

The right-hand column is not a straw man. Those are the numbers that appear in AI programme updates, and they share a property: they measure activity by the system rather than change in the business, so they go up whether or not the project worked.

## Why a holdout beats before-and-after

Between the baseline and the result, your business changed: seasonality, headcount, a pricing change, a process fix somebody made unrelated to this. A before-and-after comparison attributes all of it to the AI, which is flattering until someone senior asks a hard question and the number does not survive it.

A holdout removes the argument. Run the new system on half the queue, or one region, or one team, and keep the other on the old process for a few weeks. It costs a little speed and it converts your result from a claim into evidence — which matters most when the result is good, because that is the case you will want to spend money on again.

## The costs proposals leave out

- Inference at real volume, rather than at the volume used in the demo.
- Human review time, where the design keeps a person in the loop — often the largest running cost and frequently uncounted.
- Evaluation and maintenance: someone has to own the system, re-run the evaluation set, and respond when a provider updates a model underneath you.
- Integration upkeep as the surrounding systems change.
- The cost of being wrong — the rework, the escalation, or the customer consequence of an error the design allows through.

## When ROI is the wrong frame

Some AI work is bought to reduce a risk, satisfy a regulator, or find out whether something is possible at all. Forcing a payback number onto that kind of work produces a fiction that everyone in the room knows is a fiction. Say plainly which category a project is in; a clearly-labelled option cost is more defensible than an invented return.

## How Bitfumes handles this

The AI Opportunity Assessment produces the arithmetic before the build, not after: a named metric per use case, the current baseline where your systems can supply one, the projected cost including inference, and an explicit list of the use cases we think will not pay back. That last list is the part that saves the most money, and it is the reason the assessment is priced low enough to be worth running even when the answer turns out to be 'not yet'.

## Frequently asked

### What is a realistic payback period for an AI project?

For back-office automation — document intake, support triage, invoice matching, internal search — a well-scoped use case commonly targets payback within four to nine months. Anything projected beyond twelve months is usually too broad, and narrowing the scope is a better response than extending the horizon.

### We have no baseline. What now?

Spend two weeks measuring before you build. It feels like a delay and it is the difference between a result and an anecdote — and in most cases the measurement itself surfaces where the real cost sits, which sometimes changes what you build.

### How do we value time saved that does not reduce headcount?

Value it as capacity redeployed, and say what it was redeployed to. Time saved that nobody can name a use for is a soft number and will be treated as one; time saved that let a team absorb 30% more volume without hiring is a hard one.

### Should the consultancy that builds it also measure it?

They should supply the instrumentation, and you should own the reporting. Insist the metric comes out of a system you already trust rather than a dashboard the vendor built, for the same reason you would not accept an auditor's self-assessment.

---

Source: https://bitfumes.com/answers/how-to-measure-ai-roi
Written by Bitfumes, an AI consultancy founded in 2017 by Sarthak Shrivastava (Docker Captain). Entry point: https://bitfumes.com/ai-assessment