Everyone Measures AI Usage. 70% Can't Measure What It Returned.
The Measurement Problem with AI Usage
Anthropic surveyed 132 of its own engineers about Claude Code. Merged pull requests per day rose 67 percent. Daily use of the tool climbed from 28 to 59 percent. Self-reported productivity gains ran between 20 and 50 percent. But then someone checked the organization's delivery dashboard and saw that the delivery metrics had not moved. That gap is the whole subject of this piece.
The Gap Between Perceived Value and Actual Results
A tool can be used constantly, rated highly by the people using it, and leave no trace on the numbers a business actually runs on. The measurement problem underneath it is bigger than one company's coding assistant. McKinsey found that 30 percent of leaders could say where the time AI freed up actually went. The other 70 percent could not. Seven out of ten organizations have people spending less time on tasks and no idea whether that turned into anything.
Usage and Tokens are Costs, Not Returns
Two numbers get reported as if they answered the ROI question, and neither does. Adoption tells you whether anyone is using the thing. It is a leading indicator and a useful one, but a tool with high adoption and no measured outcome has produced activity, not value. Token spend tells you what the tool costs to run. It belongs in the calculation, on the cost side, and watching it closely tells you nothing about whether the work it produced was worth having.
The Number That Matters
The number that matters sits one step further out and takes real work to produce: whether the company made or saved a defensible dollar. Everything below is how you get to that number without lying to yourself on the way. So when a business case lands on a desk claiming a figure in saved dollars, the honest question is not whether AI helped. It is whether the number is real.
Inflation One: Soft Hours Counted as Hard Dollars
Here is the calculation almost every AI business case runs. The tool saved each person two hours a week. Multiply the hours by the hourly cost of those people, add it up across the team, and report the total as money saved. The hours are usually real. The dollars usually are not, because the budget did not change. Nobody was let go, no contractor was dropped, no line item fell.
When Do Freed Hours Become Real Money?
Freed hours become real money in three situations: the time is redeployed onto work that generates value, or it lets you avoid a hire you were about to make, or it lets the same headcount produce more of something you sell. If none of those is true, the saving is soft, and soft savings do not survive contact with a CFO who can see the budget did not drop.
The Fix: Label Every Dollar as Hard or Soft
The fix is not complicated. Label every dollar of claimed value as hard or soft, and report the two separately. The number gets smaller, and much harder to dispute.
Inflation Two: Calendar Time Counted as Labor Time
The second inflation is subtler, and I see it most in engineering cases. A feature "took three weeks" before and "takes one week" now, so the case dollarizes two weeks of saved time at an engineer's rate. The problem is that three weeks was never three weeks of work. Some of it was a ticket sitting in a queue, some was waiting on a review, some was a dependency that had not shipped.
Cycle Time vs. Labor Time
Cycle time, the calendar span from request to delivery, is not the same as labor time, the hours a person actually spent. Converting the calendar span to dollars at an hourly rate invents labor that no one performed. Speed is still worth reporting, but as a rate: this class of work now moves through 40 percent faster. It becomes money only when moving faster captures something real, most often revenue that arrives earlier because the thing shipped sooner.
What Actually Produces Dollars
Strip out the inflations and there are only two mechanisms by which an AI tool produces money, and each converts to dollars through a different bridge. The first is acceleration: someone does a task they already did, in less time. The bridge is hours saved multiplied by the hourly cost of that person. If a task dropped from four hours to two and a half and happens eighty times a month, that is 120 hours a month, and at their hourly cost you have a real figure.
The Second Mechanism: Avoided Work
The second mechanism is avoided work: a task stops happening at all. A support ticket the knowledge base resolves is a ticket a human never touches. The bridge here is not an hourly rate, it is the full cost of one whole interaction: the total monthly cost of the function divided by the number of interactions it handles.
The Denominator Nobody Writes Down
A return needs a cost to divide by, and this is where the tokens finally belong. The total cost of an AI tool per month is the build cost amortized over the months it will run, plus token spend, plus infrastructure, plus maintenance, plus any human review of its output.
How to Avoid Double-Counting
Each metric adds its own slice of value to the same numerator, and every slice divides by that same total cost. The one discipline this requires is avoiding double-counting: if a token cost already sits inside a per-unit figure, it does not also go in the denominator, and if two metrics describe the same saved dollar from two angles, you keep one of them.
The Math is Deterministic
Measuring your own AI's return is the same shape seen from the other side. The calculation is deterministic: a subtraction, a multiplication, an hourly cost, a total. No model is required, and none should be trusted with it.
Agentic AI and ROI
But isn't agentic AI supposed to be exempt from ROI? There is a serious version of the opposite argument, and Gartner makes it: early agentic AI is experimental, and organizations that demand a proven business case before they will touch it risk being outpaced by the ones that treat it as something to iterate on.
Where to Start
Capture the baseline before you deploy anything, because once the old way is gone you cannot reconstruct how long it used to take. Measure one real unit of work end to end, from request to delivered, including the waiting, and see whether that number moved. Stop reporting seats deployed, tasks completed, prompts submitted and self-reported speed, and start reporting the one figure that ties to a customer or a budget.
Questions This Raises
- How do I put a dollar value on time saved? Multiply the hours saved by the hourly cost of the person who saved them. Count it as a real saving only if that time is redeployed, avoids a hire, or produces more sellable output. Otherwise it is a soft number and should be labelled as one.
- Should token spend count as part of ROI? Yes, as a cost, in the denominator, and never as a benefit. Track cost per unit of useful output if you want a figure to hold against value, and make sure you are not counting the same tokens in two places.
- Can one project have several ROIs? No. A project has one ROI. Several metrics can each add a slice of value to the same numerator, divided by the same total cost. If you end up with two ROIs for one tool, you have either double-counted or mixed two projects together.
Comments
No comments yet. Start the discussion.