Insight · August 26, 2026
What did it
actually return?
Almost everyone is running AI agents now. Almost no one can say what they earned back. The number is provable. Most teams just never set it up to be read.
01 · The question
The room went quiet when someone asked what it earned.
In McKinsey's 2025 global survey, 88 percent of organizations said they use AI in at least one function, up from 78 percent a year before. Adoption is close to universal. Then the same survey asked the harder question. Only 39 percent could attribute any measurable profit impact to that AI, and most of those said it moved under 5 percent of earnings.
MIT's study that year, on the state of AI in business, put a sharper edge on it. Across the organizations they looked at, 95 percent saw no measurable return on their generative AI spend. Not 5 percent short of the target. No measurable return at all, against tens of billions of dollars committed.
02 · Why the number is missing
A demo is easy to admire. A return is hard to prove.
The MIT authors were clear that the failure was not the models. It was the approach. Teams stood up a pilot, watched it do something impressive, and never wrote down what done was worth before they started. So when finance asked for the return, there was no baseline to measure against and no agreed way to read it.
Gartner saw the same wave from the vendor side. It predicted that over 40 percent of agentic AI projects would be canceled by the end of 2027, on escalating cost, unclear business value, and weak controls. It also named the noise around them, agent washing, the rebranding of ordinary chatbots and automation as agents. Of thousands of vendors claiming the label, Gartner put the count of real ones near 130.
None of that means the return is not there. It means the return has to be built into the work on purpose, the same way you would for any other line on the budget.
The one line to keep
“The agent that cannot be measured is the first one cut.”
03 · The four numbers that hold up
Ignore the dashboard. Track these four.
01
Task completion
The share of jobs the agent finishes correctly with no human touch. Not attempts, not deflections. Work that is actually done and would pass review.
02
Error rate
How often the agent is wrong, and how bad the worst wrong looks. A cheap agent that fails quietly costs more than a slow human who flags the edge case.
03
Time recovered
Hours the agent gives back to a person, priced at what that person costs. This is where most of the honest return lives.
04
Cost of the runs
Tokens, infrastructure, the build, and the oversight. Every run is a real invoice. An agent that checks itself runs the model several times per task, and that shows up here.
04 · Measure the before
You cannot prove a return without a baseline.
Before the agent touches the work, measure the work. How long the task takes a person today. How often it goes wrong. What one pass costs when a human does it. Write those down. That is your baseline, and it is the single step most teams skip on the way to a year of arguing about vibes.
The reason the skip is fatal is simple. After the agent ships, everything moves at once. Volume changes, the team changes, the process changes. Without a number from before, you can never separate what the agent did from what the quarter did. The baseline is what turns a story into a measurement.
Keep the scope small enough to read. One workflow, one queue, one clear finish line. A narrow before and after that everyone trusts beats an enterprise wide claim that no one can check.
05 · The formula, worked in the open
Return minus cost, over cost. Everything else is decoration.
Here is an illustration, not a client, with round numbers so the shape is clear. A team routes a queue of routine tickets to an agent. Baseline first. A person clears each ticket in eight minutes at a loaded cost near four dollars. The agent now handles a thousand of them a week and finishes 800 cleanly, with 200 sent to a person to check or correct.
The return is the human time the agent gave back, 800 tickets at four dollars, so 3,200 dollars a week. The cost is the model runs, the infrastructure, and the share of a person's week spent reviewing the 200 and maintaining the agent. Say that lands near 1,400 dollars. The return is 3,200, the cost is 1,400, so the net is 1,800 and the ratio is a little over two to one.
ROI = (value returned minus cost of the agent) / cost of the agent
value returned = hours recovered priced at the loaded human rate
plus revenue the agent influenced, only if traceable
cost of agent = model runs plus infrastructure plus build plus oversight
Count only work the agent finished correctly.
Price the runs at what they actually cost, retries included.
If you cannot trace a number to the baseline, leave it out.One workflow · measured before and after · nothing borrowed
06 · What breaks the number
The flattering metrics are the ones that lie.
Containment is the classic trap. It counts a ticket as handled the moment no human touched it, which includes every person who gave up and walked away. An agent can post a perfect containment rate while quietly making customers angrier. Resolution, work that actually solved the problem, is the number worth trusting.
Attribution is the second trap. When an agent, a new campaign, and a pricing change all land in the same quarter, do not hand the agent the whole lift. Trace only what the baseline lets you trace. A smaller number you can defend is worth more than a big one you cannot.
And do not forget the oversight. The reason IDC could report strong returns, near four dollars back for every dollar spent in its 2024 study, and MIT could report almost none in the same period, is that the winners counted the full cost, including the people who watch the agent, and still came out ahead. The losers never counted at all.
07 · When the honest answer is zero
Sometimes the return is negative. Say so early.
Measurement is not there to protect the project. It is there to tell you the truth in time to act on it. When the four numbers come back and the cost is beating the return, the discipline is to name it, narrow the scope, or stop. That is what most of Gartner's canceled projects will be, and the ones that read the number early will lose the least.
The upside of measuring honestly is that it also finds the real winners. MIT put the share of organizations pulling serious value from their pilots near 5 percent. Those are not the teams with the best models. They are the teams that picked one workflow, set a bar they could read, and kept the ones that cleared it. The same instinct sits under how AI is reshaping what work is worth.
Closing
A return you can read is a return you can keep.
Pick one workflow. Measure it before the agent arrives. Track the four numbers, price the runs honestly, and read the result out loud. Do that once and you will never again be the team that spent the money and could not say what it bought.
Sources · McKinsey, The State of AI, 2025 · MIT, The GenAI Divide, State of AI in Business, 2025 · Gartner press release, June 2025 · IDC, 2024 Business Opportunity of AI, sponsored by Microsoft
Share this perspective
More insights
Adjacent perspectives.
Bttr. Field Brief
The brief Bttr. writes for senior buyers.
Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.