Ask ten executives what their AI ROI is, and you’ll get two liars, seven shrugs, and one person quietly changing the subject to “learnings.” The research backs up the shrugs. MIT found 95% of enterprise GenAI pilots deliver zero measurable P&L impact. BCG puts the share of companies generating meaningful financial value from AI at about a quarter; Morgan Stanley found only 21% of the S&P 500 could cite a measurable AI benefit of any kind. Meanwhile, the spend keeps climbing, because nobody wants to be the one who turned off the token fountain.
So before anyone claims that ROI is the answer, it’s worth asking a question that sounds insultingly basic: Just what is ROI, actually? There is a whole history underneath the acronym, with a birthday, an inventor, and a lesson everyone chasing “AI ROI” has likely never heard.
The gunpowder metric
Return on investment was invented in 1914 by a 29-year-old engineer named Donaldson Brown, assistant treasurer of the DuPont company, which at the time made its money on explosives. DuPont had a problem modern CIOs would recognize: capital pouring into wildly different operations (gunpowder plants, chemical works, and soon a struggling car company called General Motors) with no way to compare them. Brown’s answer was a single ratio: earnings against the capital invested to produce them.
But here’s the part that matters, the part that made ROI conquer American industry by the 1950s: Brown’s real invention sat a layer below the ratio, in the decomposition. His DuPont formula broke ROI into a tree of drivers: profit margin times asset turnover, each of those splitting further into things a plant manager could touch. The genius lay in connecting the boardroom number to the operational levers, so that when the ratio disappointed, you knew which lever was the problem. ROI worked because every term in it was measurable, and every measurement pointed to an owner.
Now look at what we call “AI ROI” today and notice what’s missing. We kept Brown’s ratio and threw away his tree.
The equation
Here’s the whole thing. AI ROI, for any workflow, is:
AI ROI = (Value created − Cost to create it) / Cost to create it
Per period, because ROI is a rate, not a trophy. Nothing exotic. A CFO from 1925 would recognize it. The reason 95% of AI projects can’t produce this number is simple: neither one survives contact with measurement. So decompose it, Brown-style, into the three questions that make it calculable:
1. Cost: What did this workflow consume? Tokens, compute, licenses, and the engineering time wrapped around them, attributed to the workflow rather than smeared across a departmental line item. This is where most companies fail before they start: enterprises report allocating 30–36% of cloud budgets to AI, while the AI spend they can trace on their bills sits closer to 2.5%. You cannot compute a ratio whose denominator you can’t find. Cost attribution is unglamorous and also the entire foundation; Brown would have started here, too.
2. Value: what changed, against what baseline? Pick the unit the workflow produces (tickets resolved, PRs shipped, claims processed, drafts delivered) and price the delta against what the same outcomes cost before. Count only the runs that succeeded, and leave the failures in the denominator; you paid for those too. A workflow that works eight times in ten at a dollar a run costs $1.25 per success, not a dollar; the failure rate is a cost multiplier hiding inside the success metric. “Engineers feel faster” doesn’t count. Measured throughput counts, along with measured quality, measured headcount-hours redeployed, and measured revenue attached. If the honest answer is “we can’t isolate the delta,” you’re holding a hope with a subscription fee. Deloitte found 74% of organizations want AI to grow revenue, and 20% have seen it. That 54-point gap is mostly in companies that never defined the unit.
3. The term Brown couldn’t have imagined: retention. Here’s where AI genuinely breaks the 1914 frame. A gunpowder plant’s return this quarter said little about next quarter’s plant. An AI workflow is different: every cycle either leaves something behind that you still own (call it a deposit) or it evaporates on contact. Two workflows with identical spend and identical output can have wildly different real returns, because one is compounding and the other is a treadmill, and the ratio alone can’t tell them apart. So the instrument needs a second readout: your retention rate, the share of AI spend that left a deposit you’re still using a quarter later. If you track only one exotic thing, track this one. It’s also the newest term of the three and the easiest to wave hands about, so it gets its own section below.
Put together, the working version anyone can calculate for one workflow this week:
AI ROI = (priced outcomes − baseline) / attributed cost, per quarter, read alongside the retention rate of the same spend.
Two readouts, one instrument. The ratio tells you what the quarter earned. The rate tells you whether next quarter’s ratio is owned or rented. Run it in a single workflow with a countable unit. The portfolio and the strategy can wait. That’s the whole method.
Retention, without the mysticism
Retention is the term that draws the most skepticism, and it should. It’s the one Brown never wrote down, and “capability you own” sounds suspiciously like the vocabulary this essay exists to ban. So hold it to the same standard as the other two terms: a unit, a rule, and a test.
The unit is the deposit. A deposit is an artifact that outlives the run that made it: a committed eval, a published prompt or skill, a documented workflow standard, a curated dataset, a step that needed a human last month and doesn’t anymore. Notice what’s not on the list: chat logs, transcripts, the “learnings” doc from the retro, the folder of experiments nobody opens. Storage isn’t retention. If a later run or another person can’t pick it up, it isn’t a deposit; it’s exhaust.
The rule is second use. An artifact counts the moment it gets used a second time without being rebuilt. Not when it’s saved. Not when it’s demoed. Used, by a later run or by someone who didn’t write it. One rule, three audit questions per workflow per quarter: Does it exist outside a chat history and someone’s head? Has anyone reused it since? Would it survive a model swap? That last question is Nadella’s sovereignty test shrunk to the size of a single artifact, the company veteran that survives the engine change. The rule also hands the term its owner, which Brown would have insisted on: whoever runs the workflow keeps the deposit ledger. Five lines a quarter (what got standardized, where it lives, when it was last reused). If the ledger is hard to write, that isn’t a reporting problem. That’s the finding.
The test is the cost curve, in arrears. You never have to price a prompt. A workflow that’s genuinely depositing shows it in the two terms you’re already measuring: cost per success falls quarter over quarter at constant scope, or the success rate climbs, because the evals catch failures before they burn tokens and the standards stop getting renegotiated from scratch. A treadmill workflow re-buys the same capability every cycle, so its curve stays flat. Retention is the one term you verify backwards: this quarter’s deposits are next quarter’s cost curve.
Consider two support workflows, identical on paper this quarter: $10,000 each, 2,000 tickets resolved, the same ROI to the decimal. Team A committed an eval suite and published the prompt that passed it. Team B left a chat history. Next quarter, A resolves the same tickets for $7,000; the evals caught two regressions before they shipped, and nobody rebuilt the prompt from memory. B pays $10,000 again, and will every quarter, forever. The ratio saw two identical workflows. The retention rate saw a compounder and a treadmill, one quarter early.
That’s the whole term. A unit you can list, a rule you can audit, a test that shows up in numbers you already collect. Brown’s bar, cleared.
Why does this defeat the 95%?
Look back at the failure studies with Brown’s tree in hand, and the pattern snaps into focus. The pilots that fail aren’t failing on model quality; MIT’s own analysis points at integration and workflow redesign, and the abandonment wave (42% of companies killed most of their AI projects in 2025) is mostly companies that could never answer the three questions. No attributed cost, no defined unit, no retention. Their ROI came back as unmeasured rather than negative, and boards eventually treat the two as the same thing. Karp’s “tokens that create no value” is really tokens whose value nobody instrumented.
The uncomfortable corollary: a company that can’t compute this equation shouldn’t conclude its AI is failing. It should conclude it’s flying blind, which is a different emergency with a different fix. Measurement first. Verdicts second.
The 1914 lesson, restated for 2026
Donaldson Brown didn’t make DuPont’s plants more profitable by exhorting them to do so. He built the instrument that let every operator see how their levers moved the number, and then the organization optimized itself; knowing was enough a century before anyone said “observability.” That’s the job in front of every company buying AI right now, and it’s humbler than a bigger model or a braver strategy deck. Build the instrument: attributed cost, a priced unit of value, and a retention rate, per workflow, per quarter.
ROI began as a way to decide which gunpowder plant deserved more capital. Yours is the same decision with a faster clock: which loops deserve more tokens?
Up close, that’s all AI ROI is: a ratio, with a tree under it. Building the instrument is the easy half. The hard part is keeping it honest once people start managing it.
The part Brown didn’t invent
That half got solved later, on factory floors rather than in treasuries. Deming, Ohno, and Goldratt were circling the same question (how do you improve a system you can only partly see, without fooling yourself?) and their answers port onto AI spend with almost no modification, which is both reassuring and a little embarrassing.
Run this monthly, on one workflow.
Find the broken branch, not the big number. Improving a non-constraint is theatre. Your constraint is rarely the most expensive workflow; it’s the one where the tree breaks, the term you can’t fill in. Sort by which of cost, value, or retention is empty.
Baseline before you change anything. One boring period measuring the current way. Every instinct says skip it, because the improvement is obvious. Skip it, and you own an anecdote for the life of that workflow. On a line, a part that ships without its measurement is a defect, however good it looks; an unbaselined workflow is the same defect.
Change one thing, for one period. A new model, or a new prompt, or a human step removed; not all three. Change three, and you get a number you can’t attribute, which is where you came in. And measure a spread, not a run: the same workflow can vary tenfold on context length, retries, and how chatty the agent decided to be. If p50 and p95 differ by an order of magnitude, that gap is the finding rather than noise to average away, usually an unbounded retry or a context nobody trims.
Standardize it or kill it. In kaizen, an improvement that never becomes standard work quietly reverts: the shift changes, the operator leaves, the knowledge goes with them. Yours reverts the same way unless the win becomes a published skill, a committed eval, or a prompt in the shared library. That step isn’t admin bolted onto the work; it is the retention term. The win becomes a deposit, and the ledger gets its line. A workflow that improved and never got standardized didn’t earn a return; it rented one.
Then the constraint moves, because it always does, and you go again.
Track one number on the loop itself before trusting any ROI it produces: coverage, the share of AI spend sitting on a workflow with a filled-in baseline. Most companies start near zero; that’s the real meaning of the 2.5% from earlier. Watch it monthly, like a yield. Once it clears half your spend, the ratio starts telling the truth.
Toyota didn’t beat Detroit by guessing which cars would sell. It beat Detroit by shortening the time between making a change and knowing whether it worked. Do that with tokens and the portfolio sorts itself (treadmills starve, compounders get fed) and from the outside it will look like you simply had better models.
You’ll know it was just a shorter loop.
← Index