Accessibility Options:
Skip to main content

AI

The AI Coding Productivity Paradox: Why the Studies Are Measuring the Wrong Thing

Koushal Dutt15 min read

The studies claiming 10–30% productivity gains from AI coding tools are not wrong — they are just measuring the wrong thing. The real story is about what developers now attempt, not how fast they complete what they were already doing.

Contents

Disclaimer: These are my personal views only and do not represent the official positions of any employer.

Financial disclaimer: This article is for informational purposes only and is not financial advice.


Let me start by conceding the point everyone uses to end the conversation: the 10% number is real.

DX Research tracked over 400 companies across 15 months (November 2024 – February 2026) and found a median 7.76% gain in PR throughput as AI usage rose 65%. That is a large-sample, longitudinal, observational study — not a press release. Yes, DX sells engineering measurement tools and has a commercial interest in positioning PR throughput as the right metric. But the number itself is credible. METR's randomized controlled trial on experienced open-source developers found AI users took 19% longer to complete tasks — not faster. The NBER working paper studying over 100,000 GitHub developers found autonomous agents drove 180% more commits, 50% more projects, and just 30% more actual releases. (Note: two of three authors are paid Microsoft research consultants, per the paper's own disclosures.)

These studies are not wrong. They are methodologically defensible. The problem is not the methodology — it is the question they are designed to answer.


What the studies actually measured

Every major productivity study on AI coding tools has a structural constraint baked in from the start: it measures tasks developers were already doing.

The METR RCT is the clearest example. It explicitly scoped itself to "large, mature, high-quality open-source codebases" — repositories with stringent documentation, testing, and style requirements, averaging 1 million lines of code. METR itself acknowledged in the same paper: "It seems plausible or likely that AI tools are useful in many other contexts different from our setting, for example, for less experienced developers, or for developers working in an unfamiliar codebase."

The DX study tracks PR throughput across enterprise engineering teams. Its own lead researcher offered the most honest line in the whole report: "Writing code was never the bottleneck." Planning, alignment, scoping, code review, and handoffs — all remain largely untouched by AI, and all are outside the numerator in any throughput metric.

The NBER weak-link finding makes the diagnosis explicit: AI dramatically accelerates raw code generation, but the elasticity of substitution between AI and human effort is just 0.25. Translation: human bottlenecks at review, QA, integration, and release absorb the gains before they show up as shipped software. The 180% more commits translated to 50% more projects and just 30% more actual releases — showing the full attenuation at each stage. The 30% more releases figure is not a ceiling on AI's value — it is a picture of what happens when you speed up one link in a chain and leave every other link human-paced.

All of this is important and worth knowing. None of it answers the question I care about more.


What the studies structurally cannot observe

I am not a developer. Everything I have built so far has been built using AI.

That sentence puts me outside the population every one of these studies is designed to measure. The DX sample is enterprise engineering teams. METR's participants were experienced open-source contributors paid $150/hour. The NBER paper tracks GitHub developers with established commit histories. Every study, by design, measures people for whom coding was already a thing they did before AI arrived.

For me, and for a growing number of people like me, AI is not a productivity multiplier. It is an enabler of creation from zero. There is no baseline to improve. The counterfactual is not "I would have done this in five days instead of four." The counterfactual is "I would not have attempted this at all."

I have been trying to build with AI since GPT-1. For a long time it was genuinely difficult to produce anything meaningful. The underlying models were impressive in narrow ways but unreliable as collaborative builders. That changed for me around November 2025, when the combination of model capability, tool integration, knowledge bases, and agents crossed a threshold — the point where nothing feels impossible anymore. That is not hyperbole. It is the practical experience of constraints collapsing. I now have a realistic sense of where those constraints still sit, developed through years of iteration. But the ceiling has moved so far up that the relevant question is no longer "can this be done with AI?" It is almost always "how do I direct the AI to do it well?"

That shift — from "can I?" to "how do I direct this?" — is precisely what no throughput study captures.

On quality: I do not just use AI to write code faster. I use it to maintain architectural discipline I could not hold alone. Providing the right context, the right reference materials, and the right constraints to a capable AI model is an architectural exercise. The AI doesn't decide what good looks like; I do. The function I perform is closer to a systems architect and technical director than a line writer. The studies do not measure that role because it does not exist in the populations they study — established developers who already know what good looks like and were building before AI existed.

On limits: the bottleneck in my own work is not AI capability. It is my own skill in directing AI — prompt quality, context-setting, knowing when to push further and when to accept what I have. That is an honest constraint. The ceiling on what I can build is not what the model can generate; it is what I can successfully specify and review. That is a very different kind of limitation than the one the NBER paper identifies.

I do not personally experience the release-cycle bottleneck the NBER paper describes, where AI-accelerated writing runs headlong into human-paced review and QA. That may be because my work sits in a different part of the distribution — greenfield builds where I control the review process rather than enterprise pipelines with separate review, QA, and release stages.


The counter-evidence the studies leave on the table

DHH (David Heinemeier Hansson) made the point clearly on the Lex Fridman Podcast #501 in August 2026: AI fundamentally changes what work you attempt rather than how fast you complete existing work. He described Omarchy Quattro — what he characterized as the most ambitious release of his professional career in terms of scope delivered — as 100% AI-generated, shipped over three months. His framing: "a limitless ceiling on ambition." The point is not that the work took less time. The point is that work of that scope and ambition got attempted at all.

The venture capital and market data make the same argument in the language of revealed preference.

Y Combinator's Winter 2025 cohort had approximately 25% of its startups reporting codebases that were 95%+ AI-generated. YC managing partner Jared Friedman noted that these were technical founders who "a year ago would have built their own product from scratch." The observation is not that they wrote code faster — it is that the nature of what they were willing to build changed.

Base44 is a concrete case: a solo founder — experienced, having previously built a 100-person VC-backed company — launched an entire platform in early 2025, acquired by Wix six months later for $80 million. Single data point, yes — but it is a verifiable public transaction, not a survey response. The capability to do that without a team is new.

Cursor's growth from $100M to $2B ARR in 13 months — the fastest revenue growth in software company history — is the kind of market signal that is hard to explain if the tool is only delivering 10% throughput gains. That is the revealed preference of engineers and enterprises, not a survey asking whether they feel more productive.

METR's own February 2026 update is particularly telling. They redesigned their study because 30–50% of developers refused to submit tasks they thought might be assigned to the AI-disallowed condition, saying working without AI would "feel so painful." METR's attempt to run a controlled experiment was defeated by the degree to which AI had become load-bearing infrastructure for working developers. That is not a productivity-gain story. That is a workflow transformation story.


We're measuring at the wrong moment

There is a well-established pattern for how general-purpose technologies enter the economy. Brynjolfsson, Rock, and Syverson described it in their 2021 paper in the American Economic Journal: Macroeconomics: when a new GPT arrives, measured productivity shows a trough before it shows a surge. The reason is that the real gains come from intangible investments — new skills, new workflows, new organizational arrangements — that are economically real but not counted in productivity statistics. Electricity took decades to show up clearly in manufacturing output data because most factories were initially wired to run old processes rather than redesigned around the new power source.

The 10% PR throughput gain is a snapshot taken at the J-curve trough. The intangible investment phase — developers learning to direct AI, teams figuring out review workflows built for AI-assisted output, founders discovering what they can now attempt — is underway but not yet in any denominator.


What we should be measuring instead

The NBER paper's weak-link diagnosis is actually the most honest frame available. AI accelerates writing code; human systems remain the bottleneck. The policy implication the paper hints at — but stops short of stating — is that the right response to AI accelerating one link is to redesign the rest of the chain, not to conclude that AI's impact is modest.

But beyond chain redesign, there are measurements the research community has not seriously attempted:

What did you build that you would not have attempted without AI? This is the capability-expansion question. No study asks it in a rigorous way. YC's 95%-AI-generated codebase data is a partial proxy, but it captures only new companies, not the full population of work being attempted for the first time.

What is the minimum viable team size for a given class of product? If a solo developer can now ship what previously required a five-person team, the productivity gain does not show up in any individual's throughput — it shows up in the org chart not being built.

What architectural decisions are now within reach? Experienced developers are attempting architecture that was previously too expensive in time and cognitive load to explore. That is not captured in commit counts.

What is the time from problem to first working prototype? The Peng et al. 55.8% speedup on greenfield tasks is the closest the literature gets to this — and it showed the largest gains, precisely because greenfield work is where the capability expansion is sharpest. (Note: Peng et al. are GitHub/Microsoft researchers; commercially interested source.)


The honest counterarguments

I am not going to wave away the signals pointing the other direction.

GitClear's analysis of 211 million lines of code found copy/pasted code rising from 8% to 18% and refactoring plummeting from 25% to less than 10% between 2021 and 2025. More code being written does not mean better code is being shipped. The manufacturing analogy holds: if you speed up the assembly line but lower the defect threshold, throughput goes up and so does rework.

Stack Overflow's 2025 Developer Survey (self-reported survey data, n ≈ 65,000) captured 84% adoption of AI tools alongside developer trust in AI output at just 29% — down from 40% the year before. Adoption rising while trust falls is an unusual pattern. The most plausible explanation is that developers are using AI tools because the productivity floor is high enough to justify it, even though they have seen enough hallucinations and subtle bugs to stop trusting the output at face value. That is a professional adaptation, not a vote of confidence in the technology's reliability.

METR's own RCT showed that experienced developers were measurably slowed on complex, mature codebases. And they showed a 39-point gap between perceived and actual speedup — developers believed AI had helped them even when it had not. Metacognitive calibration in AI-assisted development is a real problem, not a minor footnote.

These are genuine concerns. They do not refute the capability-expansion argument, but they do set the boundary conditions: AI coding tools are not uniformly valuable, they introduce quality risks if used without architectural discipline, and the productivity narrative requires careful qualification.


The paradox is not that the gains are modest

The paradox is that the most important thing AI coding tools have changed is not visible in any of the studies that have tried to measure them.

The 10% PR throughput gain is real. So is the 19% slowdown on complex codebases. Both are true simultaneously, because they are measuring different populations doing different work. Neither number tells you what a solo builder with taste and clear intent can now accomplish in three months that would previously have required a funded team. Neither number captures the developer who refused to submit a task for a study because working without AI would "feel so painful." Neither number appears in the founding story of a company that sold for $80 million six months after it was started by one person.

DHH framed it as "a limitless ceiling on ambition." I would frame it more carefully: not limitless, but the ceiling has moved so far up that the studies still measuring throughput on tasks developers were already doing are looking at the floor.

The thing that will matter in five years is not whether AI made developers 10% faster at what they were doing. It is what developers, and people who were never developers at all, will have attempted — and shipped — because the bar for attempting it finally got low enough.


Sources / Further reading