Return on Tokens (ROT): Is Your AI Spend Actually Generating Value?

— by

Uber burned through its entire 2026 Claude Code budget by April. By May, their COO went on record saying the link between all that AI consumption and actually shipping features “is not there yet.”

That cracked something open. A consultant reported a client accidentally spent $500 million on Claude Code. Amazon shut down its internal AI leaderboard. Sam Altman admitted on CNBC that AI waste has become a “huge issue.” And Ramp’s head of product called the whole dynamic the “Token Casino” — useful software wrapped in mechanics that make spend feel like progress.

If you’re in finance or accounting and you’re being asked to approve AI budgets, sit in on AI governance conversations, or explain to a board what the organization is actually getting out of its AI spend — this is your problem now. So let’s talk about what went wrong, and what a better framework looks like.

Every Cycle Has Its Dumb Metric

The pattern isn’t new. In the railroad boom of the 1850s, the market grabbed onto track miles as a proxy for future monopoly power — so railroads raced to lay duplicative track along the same routes just to make the number go up. In the dot-com era, it was website eyeballs. In the 2010s, it was gross revenue — which is how you got WeWork reporting billions in revenue while burning cash.

Each of those metrics had a real signal in it. But when the metric becomes the goal instead of the proxy, behavior gets distorted fast.

AI’s version of this is tokens. When the major labs shifted from flat subscription pricing to consumption-based billing in 2025, token spend became both how labs made money and how companies internally reported AI success. Employees who directed the most AI activity got called “AI Innovators.” Boards measured progress by how much token spend went up. The term that emerged for this behavior is tokenmaxxing — maximizing token consumption as a goal in itself, decoupled from business outcomes.

Token consumption on developer platforms went from near zero in early 2025 to over 12 trillion tokens per month by mid-2026. The question nobody was asking loudly enough: what did we actually get for it?

The Return on Tokens Framework

The fix starts with asking the right question. Not “how many tokens did we spend?” but “what did we get back?”

The formula is straightforward:

ROT = (Value of Output − Cost of Tokens) / Cost of Tokens × 100

Two levers: increase the value of what you produce, or spend less to produce it. Ideally both. Most companies right now are only pulling the cost lever — routing work to cheaper models. That helps. But it’s still the wrong frame if you’re sending AI agents to do work that deterministic code handles better, faster, and more reliably.

Why Agents Produce Negative ROT for Most Enterprise Work

Here’s a real example that stuck with me. A team used an AI agent to classify records. It was hitting 80–85% accuracy. Sounds useful, right? Except the team couldn’t tell which 80% was correct. So every single record still had to be manually reviewed. The AI was running the whole time, spending tokens, and saving almost no time at all.

This is the core problem with agents for high-stakes enterprise work. Agents improvise — they approach each task fresh, without the consistency that production-grade processes require. For fraud detection or underwriting decisions, 80% accuracy isn’t a partial win. It’s a 0% usable result.

There are three structural reasons agents struggle here. First, accuracy often doesn’t reach the threshold required for the output to be trusted. Second, most process knowledge lives in people’s heads as tacit rules, not in documentation that agents can ingest. And third, without a clear, measurable objective, agent output decays into expensive noise with no feedback loop to improve it.

The right division of labor: AI agents for judgment-heavy, ambiguous, natural-language tasks where humans would otherwise spend expensive time reasoning. Deterministic code for repetitive, rule-bound work that needs to run reliably at scale. The mistake has been using agents for both.

What This Means for Finance and Accounting Professionals

Two areas where this lands directly in your lap.

Budget and spend approval. When AI teams report progress, the default justification is activity — token consumption up, agents deployed, code shipped. Those are input metrics. When you’re approving AI budgets, push for output metrics instead: What’s the accuracy rate? What’s the cost per output unit versus the manual process? What cost reduction or revenue impact is this generating? If the team can’t answer those questions, you have a spend approval problem.

Internal controls and governance. AI token spend without output benchmarks is an unaudited cost center. The spend is real, growing, and easy to measure. The return is often vague and unquantified. That gap is a controls risk. Finance professionals embedded in AI governance conversations should be treating output measurement as a controls issue — requiring an accuracy baseline, a cost-per-outcome metric, and a benchmark against the prior process before approving ongoing spend.

This is measurement discipline. It’s something finance people are already good at. It’s exactly what’s been missing from most AI deployments.

Key Takeaways

  • Token spend is not a proxy for AI value. Reject any reporting framework that treats consumption as a success metric.
  • Apply ROT like you’d apply ROI. Every AI investment needs a measurable output — accuracy rate, time saved, cost reduced, revenue generated.
  • Ask about architecture, not just cost. Routing to cheaper models helps, but the bigger question is whether AI is being used for the right work at all.
  • Treat output measurement as a governance issue. Token spend without output benchmarks is an unaudited cost center. Finance should be flagging this.
  • Measurement discipline is the competitive advantage. The companies that get this right now will have a structural edge. Adoption without accountability leads to waste.

Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: [lesson link]

Today’s lesson


Leave a Reply

Discover more from EverydayCPE

Subscribe now to keep reading and get access to the full archive.

Continue reading