What the peer-reviewed literature says, ranked

Ten studies ranked on identification strength, sample size, outcome relevance, and peer-review status together. The strongest findings contradict the vendor narrative in ten specific places.

The short answer

What does the research show about What the peer-reviewed literature says, ranked?

Ten studies ranked on identification strength, sample size, outcome relevance, and peer-review status together. The strongest findings contradict the vendor narrative in ten specific places. Analyst firms are one tier. This is a different tier, and almost nobody in revenue reads it. The ranking below weighs identification strength, sample size, outcome relevance, and peer-review status together, so a 16-developer randomized study can outrank a 25,000-respondent survey.

Evidence

  • AI makes your best people better is the opposite of what is measured. The lowest skill quintile gained 36 percent; the most skilled got measurably worse on resolution rate and customer satisfaction.
  • The 30 to 55 percent productivity gains come from short single-task lab experiments. The 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent CI of 21 to 89 percent.
  • Individual gains do not aggregate. Economy-wide self-reported time savings equal 1.4 percent of total work hours, and a 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years.

Supporting pages

Analyst firms are one tier. This is a different tier, and almost nobody in revenue reads it. The ranking below weighs identification strength, sample size, outcome relevance, and peer-review status together, so a 16-developer randomized study can outrank a 25,000-respondent survey.

Two findings sit at the centre. The Quarterly Journal of Economics support study, 5,172 agents and roughly three million chats, found the bottom skill quintile gained 36 percent while the most skilled saw no significant productivity change and showed small but statistically significant decreases in resolution rate and customer satisfaction. And METR's developers forecast a 24 percent speedup, delivered a 19 percent slowdown, and still believed afterward that AI had sped them up by 20 percent.

That 43-point gap is the control condition for almost every AI productivity number in circulation, because almost all of them are self-reported time savings.

What this page is

Ten studies from the peer-reviewed and pre-registered literature, ranked on identification strength, sample size, outcome relevance, and peer-review status together.

The argument

The strongest identification in the literature contradicts the vendor narrative in ten specific places, and the most important contradiction is that self-reported productivity gains are unreliable in a measurable direction.

How to read it

  • The ranking weighs identification strength alongside sample size, so a 16-person randomized study can outrank a 25,000-respondent survey.
  • Only three of the ten strongest studies are peer reviewed in a named journal. The rest are working papers or pre-registered experiments.
  • Effect sizes from short single-task lab experiments do not transfer to multi-week knowledge work. The confidence intervals say so directly.

Only three of the ten strongest studies are peer reviewed in a named journal.

Ranked on identification strength, sample size, outcome relevance, and peer-review status together. Bar length encodes sample size on a log scale.

Sample size, log scale

  • 1. Brynjolfsson, Li and Raymond, QJE 2025. 5,172 support agents, about 3 million chats3.71

    Peer reviewed. Embedded RCT plus instrumental variable, objective outcome, documented skill gradient.

  • 2. Humlum and Vestergaard, BFI 2025. About 25,000 workers, Danish administrative registers4.4

    Working paper. Outcomes from tax and payroll records. Intervals tight enough to rule out effects above 1 to 2 percent.

  • 3. Dillon, Jaffe, Immorlica and Stanton. 66 firms, 7,137 workers3.85

    Preprint. Largest firm-randomized generative AI experiment. Gains confined to time reallocation.

  • 4. Dell'Acqua et al., Organization Science 2026. 758 BCG consultants2.88

    Peer reviewed, preregistered. Outside-frontier task produced a significant negative effect of minus 0.245, SE 0.054, p < 0.01.

  • 5. Fang, Yuan, Zhang, Donati and Sarvary. Seven e-commerce field experiments4.7

    Preprint. Up to 13 million subjects per experiment, revenue as the outcome.

  • 6. Otis et al. 640 Kenyan entrepreneurs2.81

    Five-month field experiment, 97 percent panel completion, pre-registered index outcome.

  • 7. Budzyn et al., Lancet Gastro Hep 2025. 19 endoscopists1.28

    Peer reviewed with a hard clinical endpoint. Observational, acknowledged residual confounding.

  • 8. METR. 16 developers, 246 real issues1.2

    Preprint. Within-developer randomization, 20 candidate explanations tested, slowdown robust from 14 to 42 percent.

  • 9. tau2-bench pass^k table0.9

    Preprint. Not a study of workers, but the most usable reliability numbers available.

  • 10. Meta-analytic central tendency, Maier et al. and Coupe and Wu2

    Preprint and working paper. g = 0.33, 95 percent CI 0.09 to 0.58 for programming, null g = 0.14 for learning.

    Deliberately ranked low

What this does not say

A low rank is not a verdict on the finding. It rates how much weight the design can carry.

Publisher
Ranked by The Revenue AI Report
Sample and method
Kosmyna EEG and the Anthropic Economic Index are deliberately ranked low: one is n = 54 with 18 in the critical session and not peer-reviewed, the other is a self-selected sample with self-reported productivity and no causal identification
Field dates
Not published by the source.
Source
No primary URL reachable at research time.

The ranking weights are ours and are stated above the chart. The sample sizes and peer-review statuses are as published.

Medium confidence
Only three of the ten strongest studies are peer reviewed in a named journal.

Ranked on identification strength, sample size, outcome relevance, and peer-review status together. Bar length encodes sample size on a log scale.

Only three of the ten strongest studies are peer reviewed in a named journal.
Sample size, log scaleValueNote
1. Brynjolfsson, Li and Raymond, QJE 2025. 5,172 support agents, about 3 million chats3.71Peer reviewed. Embedded RCT plus instrumental variable, objective outcome, documented skill gradient.
2. Humlum and Vestergaard, BFI 2025. About 25,000 workers, Danish administrative registers4.4Working paper. Outcomes from tax and payroll records. Intervals tight enough to rule out effects above 1 to 2 percent.
3. Dillon, Jaffe, Immorlica and Stanton. 66 firms, 7,137 workers3.85Preprint. Largest firm-randomized generative AI experiment. Gains confined to time reallocation.
4. Dell'Acqua et al., Organization Science 2026. 758 BCG consultants2.88Peer reviewed, preregistered. Outside-frontier task produced a significant negative effect of minus 0.245, SE 0.054, p < 0.01.
5. Fang, Yuan, Zhang, Donati and Sarvary. Seven e-commerce field experiments4.7Preprint. Up to 13 million subjects per experiment, revenue as the outcome.
6. Otis et al. 640 Kenyan entrepreneurs2.81Five-month field experiment, 97 percent panel completion, pre-registered index outcome.
7. Budzyn et al., Lancet Gastro Hep 2025. 19 endoscopists1.28Peer reviewed with a hard clinical endpoint. Observational, acknowledged residual confounding.
8. METR. 16 developers, 246 real issues1.2Preprint. Within-developer randomization, 20 candidate explanations tested, slowdown robust from 14 to 42 percent.
9. tau2-bench pass^k table0.9Preprint. Not a study of workers, but the most usable reliability numbers available.
10. Meta-analytic central tendency, Maier et al. and Coupe and Wu2Preprint and working paper. g = 0.33, 95 percent CI 0.09 to 0.58 for programming, null g = 0.14 for learning.

Source: Ranked by The Revenue AI Report. Kosmyna EEG and the Anthropic Economic Index are deliberately ranked low: one is n = 54 with 18 in the critical session and not peer-reviewed, the other is a self-selected sample with self-reported productivity and no causal identification Confidence: Medium.

Caveat: The ranking weights are ours and are stated above the chart. The sample sizes and peer-review statuses are as published.

What this does not say: A low rank is not a verdict on the finding. It rates how much weight the design can carry.

Sixteen experienced developers forecast a 24 percent speedup and delivered a 19 percent slowdown.

Predicted against measured effect on task completion time. A 43-point gap between belief and outcome.

Percent change in speed, positive is faster

  • Economics experts predicted39%
  • Machine-learning experts predicted38%
  • Developers forecast for themselves24%
  • Developers still believed afterward20%

    Measured after the trial had already shown a slowdown

  • Developers actually measured-19%

    43-point gap between the forecast and the outcome

Top to bottom spread: 43 points between belief and outcome

What this does not say

Sixteen developers is not the industry. It is the cleanest identification available on this question.

Publisher
METR, arXiv 2507.09089
Sample and method
n = 16 developers, 246 real issues in repositories they had worked in for years, within-developer randomization with screen recordings
Field dates
Not published by the source.

The clustered confidence interval on the headline estimate is not reported in the released paper.

High confidence
Sixteen experienced developers forecast a 24 percent speedup and delivered a 19 percent slowdown.

Predicted against measured effect on task completion time. A 43-point gap between belief and outcome.

Sixteen experienced developers forecast a 24 percent speedup and delivered a 19 percent slowdown.
Percent change in speed, positive is fasterValue (%)Note
Economics experts predicted39
Machine-learning experts predicted38
Developers forecast for themselves24
Developers still believed afterward20Measured after the trial had already shown a slowdown
Developers actually measured-1943-point gap between the forecast and the outcome

Source: METR, arXiv 2507.09089. n = 16 developers, 246 real issues in repositories they had worked in for years, within-developer randomization with screen recordings Confidence: High.

Caveat: The clustered confidence interval on the headline estimate is not reported in the released paper.

What this does not say: Sixteen developers is not the industry. It is the cleanest identification available on this question.

AI raised the floor by 36 percent and did nothing for the ceiling.

Productivity effect by skill quintile among 5,172 support agents.

Percent change in issues resolved per hour

  • Bottom quintile36%
  • Second quintile21%
  • Third quintile14%
  • Fourth quintile7%
  • Top quintile0%

    No significant productivity change, plus small but statistically significant decreases in resolution rate and customer satisfaction.

What this does not say

A larger gain for the bottom quintile is not evidence the top quintile is harmed. Their measured gain is smaller, not negative.

Publisher
Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025
Sample and method
5,172 support agents, approximately 3 million chats. Average effect 15 percent, bottom quintile 36 percent
Field dates
Not published by the source.

Intermediate quintile values are interpolated from the published gradient. The endpoints, 36 percent and no significant change, are as published.

High confidence
AI raised the floor by 36 percent and did nothing for the ceiling.

Productivity effect by skill quintile among 5,172 support agents.

AI raised the floor by 36 percent and did nothing for the ceiling.
Percent change in issues resolved per hourValue (%)Note
Bottom quintile36
Second quintile21
Third quintile14
Fourth quintile7
Top quintile0No significant productivity change, plus small but statistically significant decreases in resolution rate and customer satisfaction.

Source: Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025. 5,172 support agents, approximately 3 million chats. Average effect 15 percent, bottom quintile 36 percent Confidence: High.

Caveat: Intermediate quintile values are interpolated from the published gradient. The endpoints, 36 percent and no significant change, are as published.

What this does not say: A larger gain for the bottom quintile is not evidence the top quintile is harmed. Their measured gain is smaller, not negative.

The two promotional workflows are the two that failed.

Seven revenue experiments with consumer-level randomization and revenue as the outcome. Non-significant results are shown at reduced weight and tagged.

  • Pre-sale chatservice and content+16.3%
  • Search query refinementservice and content+2.9%
  • Product descriptionsservice and content+2.1%
  • Marketing push messagespromotional+1.6%n.s.
  • Google ad title optimizationpromotional-4.5%n.s.

What this does not say

Consumer field experiments do not transfer to B2B sales cycles. The outcome measured is consumer revenue.

Publisher
Fang, Yuan, Zhang, Donati and Sarvary
Sample and method
Seven field experiments, up to 13 million subjects per experiment, consumer-level randomization, revenue as the outcome. Pre-sale chat significant at p < 0.01. Preprint
Field dates
Not published by the source.

Two of the seven experiments are not charted because the published effect is reported without a directional point estimate.

High confidence

Also in the record

Figures that sit alongside these charts.

  • AI makes your best people better is the opposite of what is measured. The lowest skill quintile gained 36 percent; the most skilled got measurably worse on resolution rate and customer satisfaction.
  • The 30 to 55 percent productivity gains come from short single-task lab experiments. The 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent CI of 21 to 89 percent.
  • Individual gains do not aggregate. Economy-wide self-reported time savings equal 1.4 percent of total work hours, and a 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years.
  • Give everyone a license and productivity follows is false in the largest firm-randomized test: only 80 percent of treated workers used the tool at all in months four through six, and the median used it in 33 percent of weeks.
  • Disclosing chatbot identity before an outbound sales call reduced purchase rates by more than 79.7 percent, even though undisclosed bots matched proficient human workers.

The brief

What is going on here, and why it matters.

The charts above are the evidence. This is the read: what the data describes, the mechanism behind it, where the argument could be wrong, and what a revenue team does about it.

01

The 43-point control condition

METR's developers forecast a 24 percent speedup, delivered a 19 percent slowdown, and still believed afterward that AI had sped them up by 20 percent. That 43-point gap between belief and measurement is the control condition for almost every AI productivity number in circulation, because almost all of them are self-reported time savings.

Apply that correction to the adoption side of the Proof Gap and the two themes lock together. The activity gains people report are exactly the class of measurement this study shows to be unreliable, in the direction of overstatement.

02

AI raises the floor and does nothing for the ceiling

The Quarterly Journal of Economics support study, 5,172 agents and roughly three million chats, found the bottom skill quintile gained 36 percent while the most skilled saw no significant productivity change and showed small but statistically significant decreases in resolution rate and customer satisfaction. AI makes your best people better is the opposite of what was measured.

That reverses the standard deployment plan. If the gain concentrates in the bottom quintile, the highest-return rollout is to new hires and struggling reps, and the highest-risk rollout is to top performers whose quality measurably declined.

03

Why individual gains do not aggregate

Economy-wide self-reported time savings equal 1.4 percent of total work hours. A 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years. The headline per-task gains of 30 to 55 percent come from short single-task lab experiments, and the widely quoted 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent confidence interval of 21 to 89 percent.

Usage explains part of the aggregation failure. In the largest firm-randomized test, only 80 percent of treated workers used the tool at all in months four through six and the median used it in 33 percent of weeks. Give everyone a license and productivity follows is false as measured.

04

The finding revenue teams should read twice

Disclosing chatbot identity before an outbound sales call reduced purchase rates by more than 79.7 percent, even though undisclosed bots matched proficient human workers. Capability was not the constraint. Disclosure was, and disclosure is increasingly not optional.

That result sits directly on top of the AI slop findings, where labeling identical human content as AI depressed engagement. Two different experiments, two different channels, the same penalty attached to perceived machine authorship.

What to do with it

The move, by seat.

CRO
Discount self-reported time savings by default. The best measurement of that gap is 43 points in the direction of overstatement.
Enablement
Target rollout at the bottom skill quintile. That is where the 36 percent gain was measured and where it did not degrade quality.
Marketing and sales
Plan the disclosure penalty into any bot-fronted motion. Disclosure cut purchase rates by more than 79.7 percent in a controlled test.

Questions this page answers

What the data says, in plain language.

What does the research show about What the peer-reviewed literature says, ranked?
Ten studies ranked on identification strength, sample size, outcome relevance, and peer-review status together. The strongest findings contradict the vendor narrative in ten specific places. Analyst firms are one tier. This is a different tier, and almost nobody in revenue reads it. The ranking below weighs identification strength, sample size, outcome relevance, and peer-review status together, so a 16-developer randomized study can outrank a 25,000-respondent survey.
What does the figure "Only three of the ten strongest studies are peer reviewed in a named journal" show?
Ranked on identification strength, sample size, outcome relevance, and peer-review status together. Bar length encodes sample size on a log scale. Source: Ranked by The Revenue AI Report. Kosmyna EEG and the Anthropic Economic Index are deliberately ranked low: one is n = 54 with 18 in the critical session and not peer-reviewed, the other is a self-selected sample with self-reported productivity and no causal identification Confidence: Medium. Caveat: The ranking weights are ours and are stated above the chart. The sample sizes and peer-review statuses are as published.
What does the figure "Sixteen experienced developers forecast a 24 percent speedup and delivered a 19 percent slowdown" show?
Predicted against measured effect on task completion time. A 43-point gap between belief and outcome. Source: METR, arXiv 2507.09089. n = 16 developers, 246 real issues in repositories they had worked in for years, within-developer randomization with screen recordings Confidence: High. Caveat: The clustered confidence interval on the headline estimate is not reported in the released paper.
What does the figure "AI raised the floor by 36 percent and did nothing for the ceiling" show?
Productivity effect by skill quintile among 5,172 support agents. Source: Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025. 5,172 support agents, approximately 3 million chats. Average effect 15 percent, bottom quintile 36 percent Confidence: High. Caveat: Intermediate quintile values are interpolated from the published gradient. The endpoints, 36 percent and no significant change, are as published.
What does the figure "The two promotional workflows are the two that failed" show?
Seven revenue experiments with consumer-level randomization and revenue as the outcome. Non-significant results are shown at reduced weight and tagged. Source: Fang, Yuan, Zhang, Donati and Sarvary. Seven field experiments, up to 13 million subjects per experiment, consumer-level randomization, revenue as the outcome. Pre-sale chat significant at p < 0.01. Preprint Confidence: High. Caveat: Two of the seven experiments are not charted because the published effect is reported without a directional point estimate.
What else sits alongside these figures?
AI makes your best people better is the opposite of what is measured. The lowest skill quintile gained 36 percent; the most skilled got measurably worse on resolution rate and customer satisfaction. The 30 to 55 percent productivity gains come from short single-task lab experiments. The 55.8 percent figure came from 95 freelancers writing one HTTP server with a 95 percent CI of 21 to 89 percent. Individual gains do not aggregate. Economy-wide self-reported time savings equal 1.4 percent of total work hours, and a 2026 multi-country survey of roughly 6,000 executives found nine in ten reporting no measurable employment impact and none on productivity at their own firms over three years. Give everyone a license and productivity follows is false in the largest firm-randomized test: only 80 percent of treated workers used the tool at all in months four through six, and the median used it in 33 percent of weeks.
Where does this data come from?
Every figure is reproduced from a named publisher: Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025, METR, arXiv 2507.09089, arXiv 2504.11436, 66-firm Copilot experiment, arXiv 2510.12049, seven e-commerce field experiments, NBER working paper 34836, NBER working paper 32487, Acemoglu. Sample, field date, and confidence are shown on each chart. Sources marked as vendor research are labelled on the page.
What could not be confirmed?
A published replication of the negativity engagement experiments on AI-specific content. None exists.

Cite this page

Permanent URL and suggested citation.

https://www.therevenueaireport.com/research/peer-reviewed-layer

Kvarfordt, Jonathan. "What the peer-reviewed literature says, ranked." The Revenue AI Report, Research Library. https://www.therevenueaireport.com/research/peer-reviewed-layer

Figures on this page are reproduced from the publishers listed below. Cite the original publisher for the underlying data, and this page for the compilation and framing.

Sources

Every publisher used on this page.

If a metric, model term, or method on this page is unfamiliar, every one of them is defined in The AI and Revenue Dictionary. Sample size, field date, and confidence tags are explained there too.

Could not confirm

What we looked for and did not find.

Claims found during research and not charted

  • A published replication of the negativity engagement experiments on AI-specific content. None exists.

How to cite this research

Written by Jonathan Kvarfordt, Founder and Principal Analyst, The Revenue AI Report. Published under CC BY 4.0.

APA

Kvarfordt, J. (2026). What the peer-reviewed literature says, ranked. The Revenue AI Report. Retrieved from https://www.therevenueaireport.com/research/peer-reviewed-layer

MLA

Kvarfordt, Jonathan. "What the peer-reviewed literature says, ranked." The Revenue AI Report, 31 Aug. 2026, www.therevenueaireport.com/research/peer-reviewed-layer.

BibTeX

@misc{kvarfordt2026peerreviewedlayer,
  author = {Kvarfordt, Jonathan},
  title = {What the peer-reviewed literature says, ranked},
  year = {2026},
  publisher = {The Revenue AI Report},
  url = {https://www.therevenueaireport.com/research/peer-reviewed-layer}
}

Share this research

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Next theme

Buyers push back on AI slop, and the receipts are in

Label identical human writing as AI and engagement falls. Human-written articles drew 5.44 times the traffic. The backlash is measurable across writing, music, and images.

Subscribe

Get the weekly issue built on this data.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.