Why Most AI ROI Numbers Are Wrong (And How Finance Can Fix That)

— by

I keep running into the same conversation with clients. They’ve deployed AI tools, leadership is excited, and someone asks: “So where’s the ROI?”

Nobody has a clean answer. And the reason isn’t that AI isn’t working. It’s that most organizations are measuring it wrong.

I’ve been dealing with this firsthand — helping teams figure out how to actually quantify AI’s impact without just recycling the vendor pitch. There’s a real measurement problem here, and finance professionals are in a uniquely good position to call it out.

The Most Common Mistake

Here’s what typically happens. A company runs an AI pilot. They recruit interested employees — people who’ve already been using AI tools on their own, who are excited about the technology. The pilot goes well. Productivity looks great. Leadership rolls it out to everyone and expects the same results.

It doesn’t happen. And the reason is selection bias. The pilot group wasn’t representative — it was a group of super users. You measured their performance and assumed it would generalize. It doesn’t.

This is what I call the attribution trap. Heavy AI users are often the ones who were already highly productive. When you compare them to lower users and attribute the gap to AI, you’re confusing correlation with causation.

Goodhart’s Law Is Lurking

There’s a second problem that kicks in the moment leadership sets AI adoption as a target. It’s Goodhart’s Law: when a measure becomes a target, it ceases to be a good measure.

The second your team knows you’re tracking AI usage rates, they’ll optimize for AI usage rates. You’ll see more visible AI activity, more AI-attributed tasks, more output volume. What you won’t necessarily see is a corresponding improvement in quality or actual business outcomes. The metric gets gamed — not maliciously, just naturally. That’s how people respond to targets.

Finance should recognize this immediately. We apply this skepticism to every other KPI. AI dashboards deserve the same treatment.

What the Data Actually Shows

A DX longitudinal study tracked 400 companies from late 2024 through early 2026. AI usage across those teams grew by 65%. Actual output — measured by pull requests shipped, a more reliable signal than raw lines of code — grew by about 10%.

Ten percent. Not 10x. That’s the real number, and it’s still meaningful. But it’s a very different story than what most AI vendors are pitching.

A METR randomized controlled trial from 2025 found something even more striking. Experienced developers predicted AI would make them 24% faster. After the study, they reported feeling 20% faster. The actual measured result? They were 19% slower. The developers felt like they were moving faster because they were doing less active work — but review cycles, back-and-forth, and course-correcting AI output ate up more time than the tools saved.

This is why self-reported surveys are such a weak measurement tool. People overestimate. It’s not dishonesty — it’s human nature. When you extrapolate “I saved about an hour a day” across a team, you end up with numbers that don’t survive scrutiny.

A Better Framework: Two Buckets

When I’m working with clients on this, I split AI ROI into two buckets: amplification and augmentation.

Amplification is about human productivity. Are your people shipping more, faster? Are they spending less time on low-value work? Are developer experience scores improving? These are the gains from AI making your existing team more effective.

Augmentation is about capacity extension. To what extent is AI doing work that a human would have had to do? You can estimate this in human-equivalent hours — how long would a person have taken to do what the agent just did? Divide that by what the agent cost you, and you get an effective agent hourly rate. If that rate is low, it’s a high-ROI deployment.

This framing works well in executive conversations. “We’re amplifying our existing team” is a different value proposition than “we’re extending our capacity without adding headcount.” Both are real. They just require different evidence.

The Finance Angle

Two things matter here specifically for finance professionals.

First, ROI reporting risk. If you’re building an AI business case for the board, the methodology behind your numbers matters. Presenting a correlation as a causal gain is a credibility problem waiting to happen. When the ROI doesn’t materialize the way the model predicted, someone will ask how the forecast was built. You want a defensible answer.

Second, KPI design. If someone in your organization is proposing to track AI productivity with a single adoption-rate metric, push back. Pair it with an outcome metric. Adoption without outcomes is just Goodhart’s Law in action.

The same instinct that makes a good finance professional skeptical of a too-clean revenue forecast applies here. AI ROI claims are often too clean. The real picture is messier — and more honest.

Key Takeaways

  • The attribution trap: heavy AI users were often already the most productive — don’t mistake that correlation for causation
  • Goodhart’s Law applies directly to AI KPIs — setting adoption as a target degrades the metric
  • Real-world data shows ~10% output gains from 65% AI usage growth — meaningful, but not 10x
  • Self-reported time savings systematically overstate actual gains — controlled trials tell a different story
  • Split AI ROI into amplification (human productivity) and augmentation (agent capacity) for cleaner analysis
  • Finance’s role: apply the same skepticism to AI dashboards as to any other KPI report

Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: [lesson link]

Today’s lesson


Leave a Reply

Discover more from EverydayCPE

Subscribe now to keep reading and get access to the full archive.

Continue reading