The AI Output Gap

— by

A peer-reviewed study out of MIT and Wharton tracked more than 100,000 GitHub developers across three generations of AI coding tools — from 2022 through early 2026. They didn’t just measure what developers produced. They followed that output all the way through the production chain to see how much of it actually shipped. And then they went one step further and checked whether the things that shipped actually got used.

Here’s the finding that stopped me: autonomous AI coding agents increased lines of code written by roughly 1,700%. Actual software releases — finished product out the door — increased by 30%. And when they looked at four major app marketplaces to see if more software was reaching and being used by end users, the answer was no. More apps being created. Zero increase in apps being consumed.

That gap — between 1,700% more input and 30% more output — has a name. And it matters a lot for how you evaluate AI claims in your practice.

Why the Gap Exists: The O-Ring Problem

Software doesn’t go from idea to product in one step. It moves through a hierarchy: lines of code become files, files get committed, commits go into pull requests where humans review them, pull requests roll up into projects, and projects become releases. At every stage above the code-writing layer, a human has to do something — review, approve, merge, sign off.

The researchers invoke what economists call the O-ring principle, named after the Challenger disaster. One failed O-ring in a long chain of components brought down the whole shuttle. The principle: automating one stage of a production process moves the bottleneck to the next human-dependent stage. It doesn’t remove the bottleneck. The ceiling on total output is set by the weakest human link, not the fastest AI stage.

Think about it this way. AI just made the code-writing machine run dramatically faster. But the quality inspector at the end of the line — the developer reviewing pull requests, the project lead approving merges — has the same number of hours they always had. You don’t get 17x more finished software. You get a backlog in front of the inspector.

What the Numbers Actually Show

The study tracked three generations of tools against the full production hierarchy — not just one metric.

Autocomplete tools like early GitHub Copilot (2022): lines of code up 228%, commits up 36%, releases up 10%. Sync agents like Claude Code (early 2025): lines of code up 741%, commits up 109%, releases up 20%. Async agents like GitHub’s coding agent (mid-2025): lines of code up roughly 1,700%, commits up 180%, releases up 30%.

Every generation is more powerful. Every generation shows the same compression pattern: enormous gains at the bottom of the stack, much smaller gains at the top. The researchers estimated the elasticity of substitution between AI output and human review effort at 0.25 — well below 1.0. That number tells you AI and human review are complements, not substitutes. More AI output upstream doesn’t free up human review capacity — it increases demand for it.

They also ran a marketplace check. New iOS apps went from roughly 40,000 per month to 100,000 per month across 2025. Total user downloads across those cohorts: flat. The share of new apps failing to reach even a minimal audience rose from about 79% to 86%. More supply. Zero demand response. Robert Solow wrote in 1987 that you could see the computer age everywhere except in the productivity statistics. The authors cite him directly — because forty years later, the pattern is back.

What This Means for Accountants

The single most useful thing a CPA takes from this study is a question to ask when evaluating any AI productivity claim — from a client, a vendor pitch, or your own firm’s leadership:

“Which layer of the production process are you measuring, and what’s the human review burden between that layer and the final delivered output?”

A client says AI made their finance team 3x more productive. Productive at which layer? If it’s draft analyses, what does the review and approval cycle look like downstream? A vendor says their tool reduces close time by 40%. Which steps? Does that include controller review and CFO sign-off? I had a conversation recently with someone evaluating an AI product — the vendor was boasting about reduction in hours. My question was simple: what happened to the accuracy rate? What happened to the approval steps? Is the time actually going away, or is it moving somewhere else?

A managing partner says the firm is shipping twice as much with AI. Is that measured at task completion or at delivered-and-accepted-by-client? Because this study says the gap between those two numbers is consistently large — and it grows larger as AI tools become more powerful, because more output gets generated faster than human review processes can absorb.

Key Takeaways

  • MIT and Wharton studied 100,000+ developers across three AI tool generations. Task-level gains were up to 1,700%. Final output gains were 30%. The gap is the story.
  • The weak-link principle explains it: AI and human review are complements. More AI input upstream increases demand on human review downstream — the bottleneck moves, it doesn’t disappear.
  • App store data confirms this reaches all the way to end users — more apps created in 2025, no increase in apps consumed.
  • The one question for any AI productivity claim: which layer is being measured, and what’s the human review burden between that layer and the final delivered output?

Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: [lesson link]

Today’s lesson


Leave a Reply

Discover more from EverydayCPE

Subscribe now to keep reading and get access to the full archive.

Continue reading