What the peer-reviewed literature says, ranked
Ten studies ranked on identification strength, sample size, outcome relevance, and peer-review status together. The strongest findings contradict the vendor narrative in ten specific places.
The short answer
What does the research show about What the peer-reviewed literature says, ranked?
Evidence
- AI makes your best people better is the opposite of what is measured. The lowest skill quintile gained 36 percent; the most skilled got measurably worse on resolution rate and customer satisfaction.
- The 30 to 55 percent productivity gains come from short single-task lab experiments. The 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent CI of 21 to 89 percent.
- Individual gains do not aggregate. Economy-wide self-reported time savings equal 1.4 percent of total work hours, and a 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years.
Supporting pages
- Dictionary plain-language definitions
Analyst firms are one tier. This is a different tier, and almost nobody in revenue reads it. The ranking below weighs identification strength, sample size, outcome relevance, and peer-review status together, so a 16-developer randomized study can outrank a 25,000-respondent survey.
Two findings sit at the centre. The Quarterly Journal of Economics support study, 5,172 agents and roughly three million chats, found the bottom skill quintile gained 36 percent while the most skilled saw no significant productivity change and showed small but statistically significant decreases in resolution rate and customer satisfaction. And METR's developers forecast a 24 percent speedup, delivered a 19 percent slowdown, and still believed afterward that AI had sped them up by 20 percent.
That 43-point gap is the control condition for almost every AI productivity number in circulation, because almost all of them are self-reported time savings.
What this page is
Ten studies from the peer-reviewed and pre-registered literature, ranked on identification strength, sample size, outcome relevance, and peer-review status together.
The argument
The strongest identification in the literature contradicts the vendor narrative in ten specific places, and the most important contradiction is that self-reported productivity gains are unreliable in a measurable direction.
How to read it
- The ranking weighs identification strength alongside sample size, so a 16-person randomized study can outrank a 25,000-respondent survey.
- Only three of the ten strongest studies are peer reviewed in a named journal. The rest are working papers or pre-registered experiments.
- Effect sizes from short single-task lab experiments do not transfer to multi-week knowledge work. The confidence intervals say so directly.
Only three of the ten strongest studies are peer reviewed in a named journal.
Ranked on identification strength, sample size, outcome relevance, and peer-review status together. Bar length encodes sample size on a log scale.
Sample size, log scale
- 1. Brynjolfsson, Li and Raymond, QJE 2025. 5,172 support agents, about 3 million chats3.71
Peer reviewed. Embedded RCT plus instrumental variable, objective outcome, documented skill gradient.
- 2. Humlum and Vestergaard, BFI 2025. About 25,000 workers, Danish administrative registers4.4
Working paper. Outcomes from tax and payroll records. Intervals tight enough to rule out effects above 1 to 2 percent.
- 3. Dillon, Jaffe, Immorlica and Stanton. 66 firms, 7,137 workers3.85
Preprint. Largest firm-randomized generative AI experiment. Gains confined to time reallocation.
- 4. Dell'Acqua et al., Organization Science 2026. 758 BCG consultants2.88
Peer reviewed, preregistered. Outside-frontier task produced a significant negative effect of minus 0.245, SE 0.054, p < 0.01.
- 5. Fang, Yuan, Zhang, Donati and Sarvary. Seven e-commerce field experiments4.7
Preprint. Up to 13 million subjects per experiment, revenue as the outcome.
- 6. Otis et al. 640 Kenyan entrepreneurs2.81
Five-month field experiment, 97 percent panel completion, pre-registered index outcome.
- 7. Budzyn et al., Lancet Gastro Hep 2025. 19 endoscopists1.28
Peer reviewed with a hard clinical endpoint. Observational, acknowledged residual confounding.
- 8. METR. 16 developers, 246 real issues1.2
Preprint. Within-developer randomization, 20 candidate explanations tested, slowdown robust from 14 to 42 percent.
- 9. tau2-bench pass^k table0.9
Preprint. Not a study of workers, but the most usable reliability numbers available.
- 10. Meta-analytic central tendency, Maier et al. and Coupe and Wu2
Preprint and working paper. g = 0.33, 95 percent CI 0.09 to 0.58 for programming, null g = 0.14 for learning.
Deliberately ranked low
What this does not say
A low rank is not a verdict on the finding. It rates how much weight the design can carry.
- Publisher
- Ranked by The Revenue AI Report
- Sample and method
- Kosmyna EEG and the Anthropic Economic Index are deliberately ranked low: one is n = 54 with 18 in the critical session and not peer-reviewed, the other is a self-selected sample with self-reported productivity and no causal identification
- Field dates
- Not published by the source.
- Source
- No primary URL reachable at research time.
The ranking weights are ours and are stated above the chart. The sample sizes and peer-review statuses are as published.
Ranked on identification strength, sample size, outcome relevance, and peer-review status together. Bar length encodes sample size on a log scale.
| Sample size, log scale | Value | Note |
|---|---|---|
| 1. Brynjolfsson, Li and Raymond, QJE 2025. 5,172 support agents, about 3 million chats | 3.71 | Peer reviewed. Embedded RCT plus instrumental variable, objective outcome, documented skill gradient. |
| 2. Humlum and Vestergaard, BFI 2025. About 25,000 workers, Danish administrative registers | 4.4 | Working paper. Outcomes from tax and payroll records. Intervals tight enough to rule out effects above 1 to 2 percent. |
| 3. Dillon, Jaffe, Immorlica and Stanton. 66 firms, 7,137 workers | 3.85 | Preprint. Largest firm-randomized generative AI experiment. Gains confined to time reallocation. |
| 4. Dell'Acqua et al., Organization Science 2026. 758 BCG consultants | 2.88 | Peer reviewed, preregistered. Outside-frontier task produced a significant negative effect of minus 0.245, SE 0.054, p < 0.01. |
| 5. Fang, Yuan, Zhang, Donati and Sarvary. Seven e-commerce field experiments | 4.7 | Preprint. Up to 13 million subjects per experiment, revenue as the outcome. |
| 6. Otis et al. 640 Kenyan entrepreneurs | 2.81 | Five-month field experiment, 97 percent panel completion, pre-registered index outcome. |
| 7. Budzyn et al., Lancet Gastro Hep 2025. 19 endoscopists | 1.28 | Peer reviewed with a hard clinical endpoint. Observational, acknowledged residual confounding. |
| 8. METR. 16 developers, 246 real issues | 1.2 | Preprint. Within-developer randomization, 20 candidate explanations tested, slowdown robust from 14 to 42 percent. |
| 9. tau2-bench pass^k table | 0.9 | Preprint. Not a study of workers, but the most usable reliability numbers available. |
| 10. Meta-analytic central tendency, Maier et al. and Coupe and Wu | 2 | Preprint and working paper. g = 0.33, 95 percent CI 0.09 to 0.58 for programming, null g = 0.14 for learning. |
Source: Ranked by The Revenue AI Report. Kosmyna EEG and the Anthropic Economic Index are deliberately ranked low: one is n = 54 with 18 in the critical session and not peer-reviewed, the other is a self-selected sample with self-reported productivity and no causal identification Confidence: Medium.
Caveat: The ranking weights are ours and are stated above the chart. The sample sizes and peer-review statuses are as published.
What this does not say: A low rank is not a verdict on the finding. It rates how much weight the design can carry.
Sixteen experienced developers forecast a 24 percent speedup and delivered a 19 percent slowdown.
Predicted against measured effect on task completion time. A 43-point gap between belief and outcome.
Percent change in speed, positive is faster
- Economics experts predicted39%
- Machine-learning experts predicted38%
- Developers forecast for themselves24%
- Developers still believed afterward20%
Measured after the trial had already shown a slowdown
- Developers actually measured-19%
43-point gap between the forecast and the outcome
Top to bottom spread: 43 points between belief and outcome
What this does not say
Sixteen developers is not the industry. It is the cleanest identification available on this question.
- Publisher
- METR, arXiv 2507.09089
- Sample and method
- n = 16 developers, 246 real issues in repositories they had worked in for years, within-developer randomization with screen recordings
- Field dates
- Not published by the source.
The clustered confidence interval on the headline estimate is not reported in the released paper.
Predicted against measured effect on task completion time. A 43-point gap between belief and outcome.
| Percent change in speed, positive is faster | Value (%) | Note |
|---|---|---|
| Economics experts predicted | 39 | |
| Machine-learning experts predicted | 38 | |
| Developers forecast for themselves | 24 | |
| Developers still believed afterward | 20 | Measured after the trial had already shown a slowdown |
| Developers actually measured | -19 | 43-point gap between the forecast and the outcome |
Source: METR, arXiv 2507.09089. n = 16 developers, 246 real issues in repositories they had worked in for years, within-developer randomization with screen recordings Confidence: High.
Caveat: The clustered confidence interval on the headline estimate is not reported in the released paper.
What this does not say: Sixteen developers is not the industry. It is the cleanest identification available on this question.
AI raised the floor by 36 percent and did nothing for the ceiling.
Productivity effect by skill quintile among 5,172 support agents.
Percent change in issues resolved per hour
- Bottom quintile36%
- Second quintile21%
- Third quintile14%
- Fourth quintile7%
- Top quintile0%
No significant productivity change, plus small but statistically significant decreases in resolution rate and customer satisfaction.
What this does not say
A larger gain for the bottom quintile is not evidence the top quintile is harmed. Their measured gain is smaller, not negative.
- Publisher
- Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025
- Sample and method
- 5,172 support agents, approximately 3 million chats. Average effect 15 percent, bottom quintile 36 percent
- Field dates
- Not published by the source.
Intermediate quintile values are interpolated from the published gradient. The endpoints, 36 percent and no significant change, are as published.
Productivity effect by skill quintile among 5,172 support agents.
| Percent change in issues resolved per hour | Value (%) | Note |
|---|---|---|
| Bottom quintile | 36 | |
| Second quintile | 21 | |
| Third quintile | 14 | |
| Fourth quintile | 7 | |
| Top quintile | 0 | No significant productivity change, plus small but statistically significant decreases in resolution rate and customer satisfaction. |
Source: Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025. 5,172 support agents, approximately 3 million chats. Average effect 15 percent, bottom quintile 36 percent Confidence: High.
Caveat: Intermediate quintile values are interpolated from the published gradient. The endpoints, 36 percent and no significant change, are as published.
What this does not say: A larger gain for the bottom quintile is not evidence the top quintile is harmed. Their measured gain is smaller, not negative.
The two promotional workflows are the two that failed.
Seven revenue experiments with consumer-level randomization and revenue as the outcome. Non-significant results are shown at reduced weight and tagged.
- Pre-sale chatservice and content+16.3%
- Search query refinementservice and content+2.9%
- Product descriptionsservice and content+2.1%
- Marketing push messagespromotional+1.6%n.s.
- Google ad title optimizationpromotional-4.5%n.s.
What this does not say
Consumer field experiments do not transfer to B2B sales cycles. The outcome measured is consumer revenue.
- Publisher
- Fang, Yuan, Zhang, Donati and Sarvary
- Sample and method
- Seven field experiments, up to 13 million subjects per experiment, consumer-level randomization, revenue as the outcome. Pre-sale chat significant at p < 0.01. Preprint
- Field dates
- Not published by the source.
Two of the seven experiments are not charted because the published effect is reported without a directional point estimate.
Also in the record
Figures that sit alongside these charts.
- AI makes your best people better is the opposite of what is measured. The lowest skill quintile gained 36 percent; the most skilled got measurably worse on resolution rate and customer satisfaction.
- The 30 to 55 percent productivity gains come from short single-task lab experiments. The 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent CI of 21 to 89 percent.
- Individual gains do not aggregate. Economy-wide self-reported time savings equal 1.4 percent of total work hours, and a 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years.
- Give everyone a license and productivity follows is false in the largest firm-randomized test: only 80 percent of treated workers used the tool at all in months four through six, and the median used it in 33 percent of weeks.
- Disclosing chatbot identity before an outbound sales call reduced purchase rates by more than 79.7 percent, even though undisclosed bots matched proficient human workers.
The brief
What is going on here, and why it matters.
The charts above are the evidence. This is the read: what the data describes, the mechanism behind it, where the argument could be wrong, and what a revenue team does about it.
The 43-point control condition
METR's developers forecast a 24 percent speedup, delivered a 19 percent slowdown, and still believed afterward that AI had sped them up by 20 percent. That 43-point gap between belief and measurement is the control condition for almost every AI productivity number in circulation, because almost all of them are self-reported time savings.
Apply that correction to the adoption side of the Proof Gap and the two themes lock together. The activity gains people report are exactly the class of measurement this study shows to be unreliable, in the direction of overstatement.
AI raises the floor and does nothing for the ceiling
The Quarterly Journal of Economics support study, 5,172 agents and roughly three million chats, found the bottom skill quintile gained 36 percent while the most skilled saw no significant productivity change and showed small but statistically significant decreases in resolution rate and customer satisfaction. AI makes your best people better is the opposite of what was measured.
That reverses the standard deployment plan. If the gain concentrates in the bottom quintile, the highest-return rollout is to new hires and struggling reps, and the highest-risk rollout is to top performers whose quality measurably declined.
Why individual gains do not aggregate
Economy-wide self-reported time savings equal 1.4 percent of total work hours. A 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years. The headline per-task gains of 30 to 55 percent come from short single-task lab experiments, and the widely quoted 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent confidence interval of 21 to 89 percent.
Usage explains part of the aggregation failure. In the largest firm-randomized test, only 80 percent of treated workers used the tool at all in months four through six and the median used it in 33 percent of weeks. Give everyone a license and productivity follows is false as measured.
The finding revenue teams should read twice
Disclosing chatbot identity before an outbound sales call reduced purchase rates by more than 79.7 percent, even though undisclosed bots matched proficient human workers. Capability was not the constraint. Disclosure was, and disclosure is increasingly not optional.
That result sits directly on top of the AI slop findings, where labeling identical human content as AI depressed engagement. Two different experiments, two different channels, the same penalty attached to perceived machine authorship.
What to do with it
The move, by seat.
- CRO
- Discount self-reported time savings by default. The best measurement of that gap is 43 points in the direction of overstatement.
- Enablement
- Target rollout at the bottom skill quintile. That is where the 36 percent gain was measured and where it did not degrade quality.
- Marketing and sales
- Plan the disclosure penalty into any bot-fronted motion. Disclosure cut purchase rates by more than 79.7 percent in a controlled test.
Questions this page answers
What the data says, in plain language.
- What does the research show about What the peer-reviewed literature says, ranked?
- Ten studies ranked on identification strength, sample size, outcome relevance, and peer-review status together. The strongest findings contradict the vendor narrative in ten specific places. Analyst firms are one tier. This is a different tier, and almost nobody in revenue reads it. The ranking below weighs identification strength, sample size, outcome relevance, and peer-review status together, so a 16-developer randomized study can outrank a 25,000-respondent survey.
- What does the figure "Only three of the ten strongest studies are peer reviewed in a named journal" show?
- Ranked on identification strength, sample size, outcome relevance, and peer-review status together. Bar length encodes sample size on a log scale. Source: Ranked by The Revenue AI Report. Kosmyna EEG and the Anthropic Economic Index are deliberately ranked low: one is n = 54 with 18 in the critical session and not peer-reviewed, the other is a self-selected sample with self-reported productivity and no causal identification Confidence: Medium. Caveat: The ranking weights are ours and are stated above the chart. The sample sizes and peer-review statuses are as published.
- What does the figure "Sixteen experienced developers forecast a 24 percent speedup and delivered a 19 percent slowdown" show?
- Predicted against measured effect on task completion time. A 43-point gap between belief and outcome. Source: METR, arXiv 2507.09089. n = 16 developers, 246 real issues in repositories they had worked in for years, within-developer randomization with screen recordings Confidence: High. Caveat: The clustered confidence interval on the headline estimate is not reported in the released paper.
- What does the figure "AI raised the floor by 36 percent and did nothing for the ceiling" show?
- Productivity effect by skill quintile among 5,172 support agents. Source: Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025. 5,172 support agents, approximately 3 million chats. Average effect 15 percent, bottom quintile 36 percent Confidence: High. Caveat: Intermediate quintile values are interpolated from the published gradient. The endpoints, 36 percent and no significant change, are as published.
- What does the figure "The two promotional workflows are the two that failed" show?
- Seven revenue experiments with consumer-level randomization and revenue as the outcome. Non-significant results are shown at reduced weight and tagged. Source: Fang, Yuan, Zhang, Donati and Sarvary. Seven field experiments, up to 13 million subjects per experiment, consumer-level randomization, revenue as the outcome. Pre-sale chat significant at p < 0.01. Preprint Confidence: High. Caveat: Two of the seven experiments are not charted because the published effect is reported without a directional point estimate.
- What else sits alongside these figures?
- AI makes your best people better is the opposite of what is measured. The lowest skill quintile gained 36 percent; the most skilled got measurably worse on resolution rate and customer satisfaction. The 30 to 55 percent productivity gains come from short single-task lab experiments. The 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent CI of 21 to 89 percent. Individual gains do not aggregate. Economy-wide self-reported time savings equal 1.4 percent of total work hours, and a 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years. Give everyone a license and productivity follows is false in the largest firm-randomized test: only 80 percent of treated workers used the tool at all in months four through six, and the median used it in 33 percent of weeks.
- Where does this data come from?
- Every figure is reproduced from a named publisher: Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025, METR, arXiv 2507.09089, arXiv 2504.11436, 66-firm Copilot experiment, arXiv 2510.12049, seven e-commerce field experiments, NBER working paper 34836, NBER working paper 32487, Acemoglu. Sample, field date, and confidence are shown on each chart. Sources marked as vendor research are labelled on the page.
- What could not be confirmed?
- A published replication of the negativity engagement experiments on AI-specific content. None exists.
Cite this page
Permanent URL and suggested citation.
https://www.therevenueaireport.com/research/peer-reviewed-layer
Kvarfordt, Jonathan. "What the peer-reviewed literature says, ranked." The Revenue AI Report, Research Library. https://www.therevenueaireport.com/research/peer-reviewed-layer
Figures on this page are reproduced from the publishers listed below. Cite the original publisher for the underlying data, and this page for the compilation and framing.
Sources
Every publisher used on this page.
If a metric, model term, or method on this page is unfamiliar, every one of them is defined in The AI and Revenue Dictionary. Sample size, field date, and confidence tags are explained there too.
Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025
5,172 support agents, approximately 3 million chats, embedded RCT and instrumental variable
https://academic.oup.com/qje/article/140/2/889/7990658METR, arXiv 2507.09089
n = 16 developers, 246 real issues, within-developer randomization
https://arxiv.org/abs/2507.09089arXiv 2504.11436, 66-firm Copilot experiment
7,137 workers, random license assignment inside real companies, objective telemetry
https://arxiv.org/abs/2504.11436arXiv 2510.12049, seven e-commerce field experiments
Up to 13 million subjects per experiment, revenue as the outcome. Preprint
https://arxiv.org/html/2510.12049v5NBER working paper 34836
Multi-country survey of roughly 6,000 executives, 2026. 69 percent use AI, nine in ten report no measurable employment impact
https://www.nber.org/papers/w34836NBER working paper 32487, Acemoglu
Calibrated TFP estimate of no more than a 0.66 percent increase over ten years
https://www.nber.org/papers/w32487Luo et al., Marketing Science
Disclosing chatbot identity before an outbound sales call reduced purchase rates by more than 79.7 percent
https://pubsonline.informs.org/journal/mksc
Could not confirm
What we looked for and did not find.
Claims found during research and not charted
- A published replication of the negativity engagement experiments on AI-specific content. None exists.
How to cite this research
Written by Jonathan Kvarfordt, Founder and Principal Analyst, The Revenue AI Report. Published under CC BY 4.0.
APA
Kvarfordt, J. (2026). What the peer-reviewed literature says, ranked. The Revenue AI Report. Retrieved from https://www.therevenueaireport.com/research/peer-reviewed-layer
MLA
Kvarfordt, Jonathan. "What the peer-reviewed literature says, ranked." The Revenue AI Report, 31 Aug. 2026, www.therevenueaireport.com/research/peer-reviewed-layer.
BibTeX
@misc{kvarfordt2026peerreviewedlayer,
author = {Kvarfordt, Jonathan},
title = {What the peer-reviewed literature says, ranked},
year = {2026},
publisher = {The Revenue AI Report},
url = {https://www.therevenueaireport.com/research/peer-reviewed-layer}
}Next theme
Buyers push back on AI slop, and the receipts are inLabel identical human writing as AI and engagement falls. Human-written articles drew 5.44 times the traffic. The backlash is measurable across writing, music, and images.
Subscribe
