Adam Hagestedt. ← all essays
Monthly essay

Time saved is not ROI

AI creates capacity. Only the workflow turns capacity into money — and most enterprises never measure the gap. A field framework for proving what your AI spend actually bought.

This continues a thread. Intelligence spend is the new cloud spend argued you should own the bill the way you own cloud. Workslop is a review tax showed how polished AI output that fails the second read quietly moves the cost onto the next person. AI at work is evolving split usage into partner-in-the-loop versus dispatch-and-walk-away. Those essays are about buying and running intelligence well. This one is about the question Finance actually asks after the invoices land: did any of it show up in the P&L? Answering it honestly means breaking two measurement habits early, one on each side of the ratio — counting saved hours as value, and counting tokens as cost.

Two studies, read side by side, explain why that question is so hard to answer honestly.

A randomized controlled trial of 758 Boston Consulting Group consultants found that on tasks inside AI's capability frontier, workers completed 12.2% more tasks, 25.1% faster, at more than 40% higher quality. A different randomized trial — METR, 16 experienced open-source developers, 246 real issues in codebases they knew intimately — found AI made them 19% slower. The same developers believed it had made them roughly 20% faster.

Both results are real. That is the entire point. There is no single number for what AI does to productivity, because the answer depends on the task, the worker, the workflow, and what happens three steps downstream. The organizations that win the next two years will stop hunting for that one number and start doing something harder.

Time saved is capacity created, not value realized. The step almost nobody measures is what the organization actually did with the capacity.

The research does not say “30% faster”

The vendor slide says AI makes knowledge workers about a third more productive. That figure is a blur of studies that measured different people doing different work at different points in the value chain. Line the good ones up and the range is enormous.

Setting What the study found What an operator should take from it
Software dev (RCT, 4,867 devs) +26% completed tasks with AI access; larger gains for less-experienced developers. Coding AI can genuinely lift output — under the right conditions.
Experienced devs (METR, 16) 19% slower on mature code they knew well — while feeling ~20% faster. Self-reported speed is not evidence. Instrument it.
Customer support (5,179 agents) +14% issues resolved per hour (+34% for novices), with better sentiment and retention. Structured, measurable work is the cleanest place to prove value.
Consulting (BCG, 758) Inside the frontier: +12–25% and +40% quality. Outside it: 19 pts less likely to be correct. Quality is not a secondary metric. It flips sign.
Knowledge work (66 firms, 7,137) ~2 fewer hours per week on email; no detectable change in task quantity or composition. Saved time did not become output on its own.
Accounting (79 firms) +55% clients served weekly; monthly close cut by 7.5 days; time shifted to higher-value work. Process metrics reveal value that “hours saved” hides.
Online retail (7 workflows) Sales effects from 0% to 16.3%, driven largely by conversion. Tie AI to an economic outcome wherever the data lets you.

Stanford's 2026 AI Index lands in the same place: gains are strongest in structured work with measurable, monitorable outputs, and much less uniform in reasoning-heavy work. So the right enterprise statement is not “AI creates a 25% productivity improvement.” It is: AI productivity is conditional. My job is to find where the frontier sits inside this company, prove it experimentally, and convert the resulting capacity into economic value.

Five things enterprises call “ROI.” Only one is.

Most AI dashboards stop measuring two levels below the thing that matters. Here is the ladder I use to locate any program — Copilot, coding agents, support agents, RAG, autonomous workflows — on a single scale.

Level Typical metric Why it is not ROI
Activity Licenses, weekly active users, prompts, tokens, acceptance rate Measures adoption, not value. “72% activated” is a receipt, not a return.
Assistance Hours saved, AI-assisted hours, human-equivalent work Better — but still modeled capacity, not captured dollars.
Task Time per task, tasks completed, per-task quality Real productivity that may never improve the wider workflow.
Workflow Cycle time, throughput, rework, exceptions, SLA Much closer to enterprise value — this is where it starts to bite.
Financial Gross profit, avoided spend, cost per transaction, loss avoided The only level a CFO ultimately signs against.

Microsoft is doing the most honest work at the assistance level. Its Agent Assisted Hours metric estimates how much conventional human work an agent displaces, using real agent telemetry plus human-time benchmarks, and it is developing a “New Assisted Hours” measure for work an agent performs in parallel. Read the fine print, though: Microsoft explicitly says these metrics do not establish output quality or direct economic outcomes on their own. Assisted hours are evidence of capacity. They are not proof of value. Nobody who sells you the tool is incentivized to say that as plainly.

Tokens are not cost, either

The value side of ROI has an activity trap: assisted hours standing in for realized value. The cost side has the same bug, and it is more dangerous because it hides inside Finance's own dashboard. Most teams meter AI cost the way they meter AI value — at the activity level. Tokens consumed, seats licensed, monthly API spend. That is the denominator's version of “72% of employees activated.” It tells you what you spent. It does not tell you what a unit of work cost.

Token price is the wrong unit because it is blind to almost everything that moves the bill. A model that is cheaper per token but needs three retries, a bigger context window, a retrieval step, and a human to check its work can cost more per completed task than a pricier model that gets it right the first time. Optimize the token and you can quietly make the task more expensive.

$4 per transaction vs a $1.20 processAn autonomous agent that costs $4 a transaction to replace a step that used to cost $1.20 is not a win because it runs unattended. It is a 3× cost increase with better latency. Autonomy is a feature, not a discount.

So price the task, not the token — then roll task cost up the same ladder you roll value up: fully loaded cost per completed task, per completed workflow, and finally the cost line a CFO recognizes. Cloud spend made this move a decade ago. Nobody mature reports cost per VM-hour anymore; they report cost per transaction and cost per customer. Intelligence spend has to grow up the same way, and faster. Meter tokens for capacity planning if you like. Never mistake them for the cost of the work.

The capacity trap

Here is the accounting mistake I see most often, dressed up as a business case. Suppose 1,000 employees cost $100 an hour fully loaded, and AI appears to save each of them three hours a week.

3 hrs × 1,000 people × $100 × 52 weeks = $15.6MThat is capacity created. It is not $15.6M of P&L improvement. Nothing has necessarily changed financially. You may simply have 1,000 people who now finish Wednesday's work at 2:00 instead of 5:00.

The $15.6M becomes economic value only when something happens to it. The company avoids hiring another 100 people. Support processes 20% more cases. Engineering ships revenue-producing features that were previously in the backlog. Sales generates more qualified pipeline. Overtime drops. Outside counsel spend disappears. Those are realizations. The saved hour is not.

So report three numbers, and never let the first quietly become the third:

Capacity created — human-equivalent work removed or added. Real, but potential.

Capacity realized — freed capacity demonstrably redirected into productive work.

Financial value realized — revenue, gross profit, avoided spend, or reduced loss actually attributable to AI.

Google's DORA team now builds ROI this way for engineering: its calculator separates tooling and training cost, the initial J-curve productivity dip, headcount reinvestment capacity, revenue from additional features, and downtime impact. It argues explicitly for reinvesting freed engineering capacity rather than reflexively converting it into headcount cuts. That is the correct instinct. Capacity you fire is a one-time saving; capacity you redeploy compounds.

AI doesn't remove bottlenecks. It moves them.

This is the finding that should reset how you scope a rollout. DORA describes AI as an amplifier of the delivery system it lands in: by 2025 it found AI associated with higher throughput and higher instability. Generate code faster and you generate larger volumes that slam into code review, testing, and deployment — the parts of the pipeline AI did not touch.

So an AI rollout that makes developers 30% faster can produce almost no enterprise value if testing, review, or release approvals remain the binding constraint. You made the fast part faster and starved the slow part with more work. This is why task-level metrics lie about workflows, and why the only measurement that matters is taken at the level of the whole process, not the keystroke.

Jagged by task, not by tool

The BCG trial is the cleanest warning here. Inside the capability frontier, consultants got dramatically better. On a task engineered to sit just outside it, AI users were 19 percentage points less likely to reach the correct answer — confidently, fluently wrong. Productivity is jagged, and it is jagged by task, not by tool. That is why a use case must never be allowed to claim success by moving up one column while wrecking another:

1/ A coding agent that generates twice the code but creates 40% more review work is not obviously productive.

2/ A support agent that lowers cost per ticket but drives repeat contacts up is not a win.

3/ A sales assistant that saves reps five hours a week but never moves pipeline has produced no financial ROI.

4/ An autonomous finance agent that saves 5,000 labor hours but introduces a material-control risk is a liability, not a success.

The framework: follow the work to the money

One structure holds all of this together. I call it the AI Value Realization Stack. It reads as a chain of evidence, not a dashboard of vanity metrics, and every layer has to be earned before you claim the next.

Cost · Quality · Risk · Human impact · Attribution confidence — measured at every layer
  1. 0 · CounterfactualWhat happens without AI? Baseline cycle time, output, quality, cost.
  2. 1 · AdoptionIs the intended population — people and agents — actually using it?
  3. 2 · TaskIs a unit of work faster or better, without hidden quality loss?
  4. 3 · WorkflowDid throughput, cycle time, or rework improve end to end?
  5. 4 · BusinessDid an outcome the business cares about actually move?
  6. 5 · FinancialDid revenue, margin, avoided spend, or expected loss actually change?

The discipline is the horizontal rail. A layer-3 throughput gain that quietly raises risk at layer 4 is not a gain. A layer-2 speed-up that degrades quality is not a speed-up. You do not get to bank one axis while breaking another.

The number a CFO will accept

At the top, the formula is boring on purpose:

AI ROI = (Realized value − Fully loaded cost) ÷ Fully loaded costThe rigor lives entirely in how tightly you define the two terms. Loose definitions are how vendor ROI decks reach 400%.

Recognize realized value in only four forms, and be strict about each: incremental gross profit (not incremental revenue); cashable cost reduction (contractors, overtime, outsourced services, retired systems); avoided future spend (chiefly hiring you did not do because capacity rose); and expected loss reduction, where risk can be honestly quantified as probability × impact. Capacity created belongs on the scorecard, but it does not get to sneak into hard-dollar ROI.

Then load the cost honestly. Not just licenses and API spend, but inference and infrastructure, integration engineering, RAG and data preparation, evaluations and observability, security, compliance, training, change management, human review, incremental QA, and the initial productivity dip. This is the task-cost ladder from earlier rolled all the way up: the honest denominator is fully loaded cost per unit of value, not a line item from the API bill. It is also where most vendor claims quietly fall apart — gross revenue in the numerator, license cost alone in the denominator. Put the real denominator in and a lot of “wins” go flat.

Put a confidence level on every dollar

Do not measure ROI by comparing power users to everyone else. Your best people were the first to become power users; you will credit AI for talent you already had. You need a counterfactual. Here is the evidence hierarchy I hold claims to, best to worst:

1/ Randomized treatment/control, or randomized phased rollout — very high confidence.

2/ Matched cohorts with difference-in-differences — high.

3/ Interrupted time series or synthetic control — medium-high.

4/ Before/after — medium-low.

5/ Users vs non-users — low. Surveyed “time saved” for financial attribution — effectively zero.

You do not always need a lab. Microsoft's 2026 Word study compared more than 72,000 sustained Copilot users against similar coworkers who adopted later, using the late adopters as a built-in control — far stronger than naive before/after. The cleanest studies simply randomized access. The practical move is to make causal confidence visible on the scorecard itself:

$24M realized · $17M causally verified · $7M directionally attributedThat sentence is more credible to a board than “$24M in AI value,” because it admits which dollars are proven and which are inferred. Honesty about attribution is a feature, not a hedge.

The macro backdrop is on your side

If this feels like more rigor than the moment demands, the aggregate numbers say otherwise. McKinsey's 2025 State of AI found nearly two-thirds of organizations had not yet scaled AI across the enterprise, and only 39% reported enterprise-level EBIT impact. Deloitte's 2025 research found most respondents expecting two to four years to satisfactory AI ROI — against the seven-to-twelve-month expectation for typical technology — with only 6% reporting payback inside a year. These are surveys, not causal proof; use them to frame the problem, not to price a use case.

Enterprises are not short on AI activity. They are short on the mechanism that converts activity into economics.

Building that mechanism is the job. It is also the differentiator, because almost nobody has built it yet.

What I'd do Monday

Start narrow, instrument first, and gate every expansion on evidence.

1/ Map workflows, not tools. List the high-labor, high-revenue, high-delay, and high-risk workflows. Decompose each into tasks and find where AI is economically promising. Do not start from an inventory of AI tools — that is how you end up measuring licenses.

2/ Open an AI Value Register. Every real use case gets a business owner, a baseline, an intervention, an eligible population, a primary KPI, quality guardrails, a fully loaded cost model expressed per completed task and workflow (not per token or seat), an experimental design, a financial hypothesis, and an attribution-confidence level. If a use case can't fill that row, it isn't ready to scale.

3/ Instrument before you scale. Capture baseline workflow telemetry first. Where you can, stagger the license or agent rollout so late adopters become your control group. You cannot reconstruct a counterfactual after the fact.

4/ Run value experiments, not demos. A prototype proves capability. A value experiment proves the operating metric moved. Fund the second, and treat a slick demo as the beginning of the question, not the answer.

5/ Gate scale on the right layer. Adoption justifies more experimentation. Task gains justify a larger pilot. A workflow improvement justifies deployment. Only financial realization justifies enterprise scaling. Do not skip a rung because the demo was impressive.

6/ Review with Finance quarterly. One page: investment, adoption, capacity created, workflow improvements, realized financial value, risk changes, and attribution confidence. Kill low-value projects aggressively and move the money toward demonstrated winners. This is a very different culture from “we bought 20,000 licenses and 74% are active.”

And resist the one metric everyone will ask you to build: a single company-wide “AI productivity: 22%.” Engineers, attorneys, recruiters, and support agents do not produce interchangeable units of output; the blended number is meaningless the moment you drill into it. Report enterprise financial value, plus portfolio-level capacity created, plus workflow-level KPI movement, plus quality and risk, plus adoption — then allow the drill-down. DORA gives engineering the same warning: never explain a complex system with one number, never compare unlike teams, and never turn a metric into a target people will game.

Don't assign a dollar value to an AI action. Follow the action through the workflow until you find the economic outcome.

That is the whole discipline. Measure the workflow, not the model. Prove the counterfactual, not the vibe. Follow the capacity until it turns into money — or admit, on the record, that it didn't. The operators who do this will be the ones who can still answer the CFO's question a year from now.

Sources

  1. Developer productivity RCTs (Microsoft, Accenture, Fortune 100), Management Science: pubsonline.informs.org
  2. METR, experienced open-source developer RCT: metr.org
  3. NBER, customer support field study (w31161): nber.org
  4. BCG / Harvard, capability-frontier study: doi.org
  5. NBER, workplace experiment across 66 firms (w33795): nber.org
  6. Stanford GSB, human–AI accounting field evidence: gsb.stanford.edu
  7. Columbia Business School, online-retail field experiments: business.columbia.edu
  8. Stanford HAI, 2026 AI Index (economy): hai.stanford.edu
  9. Microsoft, Agent Assisted Hours: microsoft.github.io · Word Copilot measurement study: microsoft.com/research
  10. Google DORA, generative-AI report and ROI calculator: dora.dev/ai · roi calculator
  11. McKinsey, State of AI: mckinsey.com · Deloitte, AI ROI research: deloitte.com