Companies have spent the last two years paying for AI coding tools on the premise that more code, written faster, means more software delivered. A new study of more than 100,000 developers suggests that premise is wrong.
A study of more than 100,000 developers finds a vast gap between writing code and shipping software. The reason is human bottlenecks

Aerps.com / Unsplash
Companies have spent the last two years paying for AI coding tools on the premise that more code, written faster, means more software delivered. A new study of more than 100,000 developers suggests that premise is wrong.
The latest AI coding tools produced 741% more lines of code. Actual software releases rose 20%. That gap — between what the tools generate and what teams actually ship — is the central finding, and it has direct implications for anyone making decisions about AI investment.
The study, from researchers at MIT and Wharton, is the first to trace the effect of AI coding tools from raw code all the way through to shipped software and real-world usage at scale. Earlier research measured how fast developers completed individual tasks. This one asks whether any of that speed shows up in the final product.
The May 2026 working paper combined public GitHub data with internal Microsoft $MSFT records. Researchers tracked developer activity at every stage of the process — from writing code to reviewing it, approving it, and finally pushing it out to users as a finished release.
The study covered three generations of tools: autocomplete, conversational assistants, and autonomous agents. Each produced larger gains in raw output than the last. Autocomplete increased the volume of code developers submitted for review by about 40%. Adding conversational assistants brought that figure to 140%. Autonomous agents raised it to 180%.
But the further along the pipeline, the smaller the gains got. Conversational assistants produced that 741% increase in lines of code written but only a 65% increase in code submitted for review and just 20% more finished releases. The tools were making developers much faster at writing code. They weren't making teams much faster at shipping it.
The researchers explain the finding through what they call the "weak-link hypothesis." Software development isn't a single task. It's a chain: writing code, integrating changes, reviewing and approving those changes, and managing releases. AI tools accelerate the first link. The later ones still require human judgment, coordination, and decision-making.
The upshot is that AI and human effort aren't substitutes at any stage beyond raw code generation. You can't replace reviewing, testing, and release management with more lines of code.
Google $GOOGL's DevOps Research and Assessment program has tracked software delivery performance for a decade and never counted lines of code among its core metrics. Code volume is an input, not an outcome.
The MIT and Wharton researchers also cross-referenced their GitHub findings with data from four major app marketplaces. More apps were being created but fewer people were using them.
Earlier studies told a simpler story. Developers using GitHub Copilot completed a single task 55.8% faster than a control group in a 2023 experiment. A later set of trials across Microsoft, Accenture $ACN, and a Fortune 100 company found a 26% increase in completed code reviews per week among nearly 5,000 developers.
Those numbers measured task completion, not the full pipeline. Doing one step faster doesn't automatically speed up everything that follows.
A separate trial by METR, a nonprofit research organization, adds a wrinkle. Developers using AI believed they finished their work 20% faster when they really took 19% longer. METR later found some evidence of speedup with newer tools but acknowledged the signal was unreliable.
The researchers acknowledge several limitations. Quality measures are indirect, relying on app-store ratings and downloads. The study covers public repositories and app marketplaces but not internal corporate software, which represents a large share of the industry.
Timing is also a factor. Code written today may ship weeks or months later, and the study can't fully resolve whether the gap is permanent or partly a matter of lag. Autonomous coding agents only became widely available in mid-2025, so the data captures early adoption. Whether teams close the gap as they adapt their workflows is still an open question.
What the data does establish is a pattern: each generation of AI coding tool widened the gap between code written and software shipped. That pattern holds at the individual developer level and at the aggregate marketplace level. And it suggests that the bottleneck in software development was never the writing of code itself.
Join 500,000+ readers who start their day with Quartz.
By subscribing, you agree to our Terms of Service and Privacy Policy.