A recent study of hundreds of software firms indicates that while AI coding agents generate more code, they do not lead to an increase in overall software output or a reduction in employment. Human code review acts as a significant bottleneck, limiting the efficiency gains provided by these automated tools.
Harvard University researchers Fiona Chen and James Stratton analyzed data from Jellyfish, a platform that tracks engineering team output. The dataset included 300 million individual work events, such as commits and pull requests, along with issue management data. It covered more than 700,000 employees at over 700 software development firms between 2021 and March 2026.
The study found that any efficiency gained during the coding phase is absorbed by downstream constraints in the production process. Specifically, the code review process becomes significantly longer when AI tools are used. Pull requests are more likely to require revisions, and reviewers leave more comments on the generated code.
Programmers who use these tools often do not trust the accuracy of the AI-generated output. Consequently, substantial effort is required to review the code. The researchers concluded that there is little evidence that firms increase software output or reduce employment by adopting AI coding agents.