Harvard study ties coding-agent gains to heavier code review
A study of 300 million work events across 718 firms found more code activity after coding-agent adoption, but no statistically significant increase in resolved issues or epics as review demands rose.
A Harvard University working paper found that firms generated more code after adopting AI coding agents, but did not record a statistically significant increase in resolved issues or epics, the study’s measures of completed software output. The analysis associates that gap with heavier downstream code review. It does not establish that coding agents caused lower code quality or reduced developer productivity overall.
The study analyzed about 300 million work events involving 725,938 workers at 718 firms from January 2021 through March 2026. Following firm-level agent adoption, lines of code added per worker-month rose by an estimated 30%, commits by 20% and pull requests by 23%, each relative to its baseline mean.
Review demands rose at the same time. Average pull-request review time, measured from submission to merge, increased by an estimated 3.45 days, or 49% from a 7.03-day baseline. The share of pull requests receiving a formal change request increased by an estimated 0.12 from a baseline of 0.13, nearly doubling. Review comments rose by an estimated 0.58 per pull request, or 35% from a baseline of 1.66.
The share of workers conducting at least one code review increased by four percentage points, or 14% from a 29% baseline. The share writing code did not change significantly.
Yet those activity gains did not produce a statistically significant change in the study’s measures of completed software. The pooled estimate for resolved issues was an increase of 0.12 per worker-month, with a standard error of 0.17, against a baseline of 3.67. The authors said the 95% confidence interval ruled out an increase larger than 12% of the baseline mean.
They also found no statistically significant effect on resolved epics and no evidence, under two issue-size measures, that the issue result reflected a shift toward larger or more complex tasks.
Fiona Chen and James Stratton used variation in firms’ adoption timing in a staggered difference-in-differences analysis, comparing changes at earlier adopters with those at later or non-adopters. The model included firm and month fixed effects, controls tied to baseline firm size and standard errors clustered by firm. Its interpretation depends on the assumption that those groups would otherwise have followed parallel outcome trends. Larger firms adopted earlier, according to the paper.
The researchers describe code review as a mechanism consistent with incomplete pass-through from code production to completed work, not definitive proof of the cause. More review could reflect a larger volume of code awaiting review, lower-quality draft code or changed review standards. The paper found no statistically significant change in lines added per pull request, which argues against larger pull requests mechanically explaining the higher comment and change-request counts.
The dataset records commit and pull-request metadata, not the code itself, so it cannot directly establish defect rates or the intrinsic quality of agent-generated code. It also measures resolved issues and epics rather than customer value, revenue, reliability or developer well-being. The proprietary data cannot be independently audited from the public paper, and the observational design cannot eliminate every unmeasured difference between early and later adopters.
Agent adoption was dated using the earliest signal from Claude Code API data or GitHub bot, commit-message and pull-request signatures. The authors said that method can miss individual licenses and firms that do not integrate their tools with Jellyfish. Detected firm-level access therefore does not imply uniform use by every worker.
More news

OpenAI reports 130,000 ChatGPT users and 95,000-plus Codex users at Oracle

Anthropic puts $100 million behind plan to train 10,000 enterprise AI engineers

OpenAI broadens Albertsons AI partnership around Safeway shopping
