The most-quoted AI numbers do not survive contact with their own methodology

How a 153-person conference survey became a Fortune 500 footnote pointing at example.com, and the twelve failure modes every number now gets run against.

The short answer

What does the research show about The most-quoted AI numbers do not survive contact with their own methodology?

How a 153-person conference survey became a Fortune 500 footnote pointing at example.com, and the twelve failure modes every number now gets run against. The claim, as it circulated: 95 percent of enterprise AI pilots fail. What is true about the document is narrower. It is not listed on the MIT Media Lab NANDA publications page, the only live copies are third-party mirrors, and the filename still says v0.1. The sample is 300-plus publicly disclosed initiatives, 52 interviews, and 153 leaders recruited at a conference, fielded January to June 2025. The listed reviewer is also a co-author. Success is operationalized as whether users or executives have remarked on impact. The ROI observation window is six months and the report's own appendix concedes that may be insufficient.

Evidence

  • The twelve house failure modes: no primary copy, a version number in the filename, reviewer is a co-author, sample built from a conference or webinar, subjective success definition, observation window shorter than the payback period, forecast presented as observation, round total numbers, underpowered for the number of tests run, attribution drift, method misstated in coverage, and a dead or placeholder citation.
  • The correctly sourced deskilling finding is stronger than the one that went viral. Endoscopists with a mean 28 years of practice saw unassisted adenoma detection fall from 28.4 percent to 22.4 percent after three months of routine AI exposure, adjusted odds ratio 0.69, 95 percent CI 0.53 to 0.89, p = 0.005.
  • The Gartner release behind the 40 percent forecast also estimated that only about 130 of the thousands of agentic AI vendors are real. That is the more useful number and it got a fraction of the coverage.

Supporting pages

The claim, as it circulated: 95 percent of enterprise AI pilots fail. What is true about the document is narrower. It is not listed on the MIT Media Lab NANDA publications page, the only live copies are third-party mirrors, and the filename still says v0.1. The sample is 300-plus publicly disclosed initiatives, 52 interviews, and 153 leaders recruited at a conference, fielded January to June 2025. The listed reviewer is also a co-author. Success is operationalized as whether users or executives have remarked on impact. The ROI observation window is six months and the report's own appendix concedes that may be insufficient.

Two people with standing said so at the time. Wharton's Kevin Werbach asked for the supporting data or a retraction. Paul Roetzer told readers not to put weight into the study. Then the citation chain decayed, and the decay is documented: Fortune described the method as 150 interviews and 350 employees, Menlo Ventures attributed it to the mirror host, and a Rockwell Automation white paper footnoted the 95 percent to example.com.

The pattern repeats. Gartner's 40 percent is a forecast, not a measurement, and the only data in the release behind it is a poll of 3,412 webinar attendees about investment posture. The Forrester 88 and 22 percent figures appear in no primary Forrester publication we could locate, so they are not charted. The half-of-entry-level-jobs claim was said in an interview with no study behind it. The properly sourced version, Stanford's Canaries in the Coal Mine, reports a 16 percent relative employment decline for 22 to 25 year olds in AI-exposed occupations, built on ADP payroll records.

What this page is

A traced citation chain showing how one conference survey became a widely quoted enterprise statistic, and the twelve failure modes now run against every number in this library.

The argument

The most-quoted AI numbers are the ones least able to survive their own methodology. Decay is not accidental, it follows a documented path, and it can be tested for.

How to read it

  • Naming a decay path is not the same as saying the underlying finding is false. It says the number as quoted is not supported by what was measured.
  • The twelve failure modes are house rules for this publication, not a validated instrument.
  • Figures that could not be located in a primary publication are not charted here, even where they are widely repeated.

A 153-person conference survey became a Fortune 500 footnote pointing at example.com.

Documentation view, four stages, one path. Every node carries its own source.

  1. 01 · The document

    v0.1 PDF, not on the MIT publications page

    n = 153 conference-recruited leaders plus 52 interviews. The listed reviewer is also a co-author.

    https://cloudelligent.com/wp-content/uploads/2026/02/v0.1_State_of_AI_in_Business_2025_Report.pdf

  2. 02 · The press

    Fortune, 18 August 2025

    Method restated as 150 interviews and 350 employees, which is not what the paper says.

    https://fortune.com/

  3. 03 · The industry

    Menlo Ventures enterprise report

    Attributed to MIT MLQ AI, which is the mirror host rather than the publisher.

    https://menlovc.com/perspective/2025-mid-year-llm-market-update/

  4. 04 · The endpoint

    Rockwell Automation white paper

    Cited the 95 percent to https://example.com/

    https://example.com/

What this does not say

A decayed citation does not make the original finding false. It makes the number being repeated unsupported.

Publisher
Chain traced by The Revenue AI Report
Sample and method
Four nodes verified, each against the document as published
Field dates
August 2026
Source
No primary URL reachable at research time.
High confidence

The most-quoted claims trip the most methodology tests.

Six widely quoted claims against six of the twelve house failure modes. The count, not the colour, is the point.

PublisherNo primary copySelf-selected sampleSubjective success definitionForecast sold as observationRound total numberAttribution driftModes tripped
MIT 95 percent of pilots fail5/6
Kosmyna brain rot2/6
Gartner 40 percent of agentic projects cancelled3/6
Forrester 88 percent and 22 percent3/6
Half of entry-level white collar jobs4/6
85 percent of AI projects fail3/6

For comparison, a number that holds

Stanford Canaries in the Coal Mine. ADP payroll records, 16 percent relative employment decline for ages 22 to 25 in AI-exposed occupations. Zero modes triggered.

What this does not say

A triggered failure mode is not proof the study is wrong. It marks where the study cannot support the weight put on it.

Publisher
Scored by The Revenue AI Report against the twelve failure modes
Sample and method
Marks show a triggered mode. Each row was scored against the primary document or, where none exists, the mirror it is served from
Field dates
Not published by the source.
Source
No primary URL reachable at research time.

Six of the twelve modes are shown here for legibility. The full twelve-mode scoring is published with each teardown.

Medium confidence

Quoted numbers land on multiples of five. Measured numbers do not.

Every gridline is a multiple of five. The upper band lands on them. The lower band does not. That is the whole argument.

Numbers everybody quotes9585805040Numbers that were measured166191420255075100
  • 95 percent of enterprise AI pilots fail, MIT NANDA v0.195
  • 85 percent of AI projects fail, traced to a removed VentureBeat article85
  • 80 percent version, appears in RAND as a hedged by some estimates80
  • Half of entry-level white collar jobs, said in an interview50
  • 40 percent of agentic projects cancelled by 2027, a Gartner forecast40
  • 16 percent relative employment decline, ages 22 to 25, Stanford on ADP payroll records16
  • 6.0 point drop in adenoma detection after AI exposure, Lancet Gastro Hep 20256
  • 19 percent developer slowdown measured against a 24 percent forecast, METR19
  • 14.41 percent WebArena best agent score against 78.24 percent for humans14
  • 2.5 percent best automation rate, Remote Labor Index2

The strongest numbers of the last two years are the ones nobody quotes, because they are narrow, hedged, and tied to administrative data. The weakest are the ones everybody quotes, because they are round and arrive with an institutional name attached to something that institution did not publish.

What this does not say

A round number is not automatically invented. It is a signal to go find the measurement behind it.

Publisher
Compiled by The Revenue AI Report
Sample and method
Each point carries its own publisher in the list below the chart. Values plotted as published
Field dates
Not published by the source.
Source
No primary URL reachable at research time.

Rounding is a signal, not a proof. A round number can be correct. It just has to show its appendix first.

Medium confidence

Also in the record

Figures that sit alongside these charts.

  • The twelve house failure modes: no primary copy, a version number in the filename, reviewer is a co-author, sample built from a conference or webinar, subjective success definition, observation window shorter than the payback period, forecast presented as observation, round total numbers, underpowered for the number of tests run, attribution drift, method misstated in coverage, and a dead or placeholder citation.
  • The correctly sourced deskilling finding is stronger than the one that went viral. Endoscopists with a mean 28 years of practice saw unassisted adenoma detection fall from 28.4 percent to 22.4 percent after three months of routine AI exposure, adjusted odds ratio 0.69, 95 percent CI 0.53 to 0.89, p = 0.005.
  • The Gartner release behind the 40 percent forecast also estimated that only about 130 of the thousands of agentic AI vendors are real. That is the more useful number and it got a fraction of the coverage.

The brief

What is going on here, and why it matters.

The charts above are the evidence. This is the read: what the data describes, the mechanism behind it, where the argument could be wrong, and what a revenue team does about it.

01

One chain, traced end to end

The claim circulated as 95 percent of enterprise AI pilots fail. The document behind it is not listed on the MIT Media Lab NANDA publications page, the only live copies are third-party mirrors, and the filename still says v0.1. The sample is 300-plus publicly disclosed initiatives, 52 interviews, and 153 leaders recruited at a conference, fielded January to June 2025. The listed reviewer is also a co-author. Success is operationalized as whether users or executives have remarked on impact, and the ROI observation window is six months, which the report's own appendix concedes may be insufficient.

Then it decayed in public. Fortune described the method as 150 interviews and 350 employees. Menlo Ventures attributed it to the mirror host. A Rockwell Automation white paper footnoted the 95 percent to example.com. Two people with standing objected at the time: Wharton's Kevin Werbach asked for the supporting data or a retraction, and Paul Roetzer told readers not to put weight into the study. Neither slowed the number down.

02

The pattern is not unique to one study

Gartner's 40 percent is a forecast presented as an observation, and the only data in the release behind it is a poll of 3,412 webinar attendees about investment posture. The Forrester 88 and 22 percent figures appear in no primary Forrester publication that could be located, so they are not charted. The half-of-entry-level-jobs claim was said in an interview with no study behind it.

The properly sourced version of that last claim is stronger than the viral one. Stanford's Canaries in the Coal Mine reports a 16 percent relative employment decline for 22 to 25 year olds in AI-exposed occupations, built on ADP payroll records. Real identification, real data, a fraction of the reach.

03

The test you can run in five minutes

The twelve house failure modes: no primary copy, a version number in the filename, reviewer is a co-author, sample built from a conference or webinar, subjective success definition, observation window shorter than the payback period, forecast presented as observation, round total numbers, underpowered for the number of tests run, attribution drift, method misstated in coverage, and a dead or placeholder citation.

Quoted numbers land on multiples of five and measured numbers do not. That single heuristic catches a surprising share of the problem before you open anything. The Gartner release behind the 40 percent forecast also estimated that only about 130 of the thousands of agentic AI vendors are real, which is the more useful number and got a fraction of the coverage.

What to do with it

The move, by seat.

Anyone building a board deck
Run the twelve tests before the number goes on a slide. A dead citation found by a board member costs more than the slide was worth.
Marketing
Link to the primary document, not to coverage of it. Attribution drift starts at the second hop.
Executives
Treat round numbers as unmeasured until proven otherwise. Measurement rarely lands on a multiple of five.

Questions this page answers

What the data says, in plain language.

What does the research show about The most-quoted AI numbers do not survive contact with their own methodology?
How a 153-person conference survey became a Fortune 500 footnote pointing at example.com, and the twelve failure modes every number now gets run against. The claim, as it circulated: 95 percent of enterprise AI pilots fail. What is true about the document is narrower. It is not listed on the MIT Media Lab NANDA publications page, the only live copies are third-party mirrors, and the filename still says v0.1. The sample is 300-plus publicly disclosed initiatives, 52 interviews, and 153 leaders recruited at a conference, fielded January to June 2025. The listed reviewer is also a co-author. Success is operationalized as whether users or executives have remarked on impact. The ROI observation window is six months and the report's own appendix concedes that may be insufficient.
What does the figure "A 153-person conference survey became a Fortune 500 footnote pointing at example.com" show?
Documentation view, four stages, one path. Every node carries its own source. Source: Chain traced by The Revenue AI Report. Four nodes verified, each against the document as published Fielded August 2026. Confidence: High.
What does the figure "The most-quoted claims trip the most methodology tests" show?
Six widely quoted claims against six of the twelve house failure modes. The count, not the colour, is the point. Source: Scored by The Revenue AI Report against the twelve failure modes. Marks show a triggered mode. Each row was scored against the primary document or, where none exists, the mirror it is served from Confidence: Medium. Caveat: Six of the twelve modes are shown here for legibility. The full twelve-mode scoring is published with each teardown.
What does the figure "Quoted numbers land on multiples of five. Measured numbers do not" show?
Every gridline is a multiple of five. The upper band lands on them. The lower band does not. That is the whole argument. Source: Compiled by The Revenue AI Report. Each point carries its own publisher in the list below the chart. Values plotted as published Confidence: Medium. Caveat: Rounding is a signal, not a proof. A round number can be correct. It just has to show its appendix first.
What else sits alongside these figures?
The twelve house failure modes: no primary copy, a version number in the filename, reviewer is a co-author, sample built from a conference or webinar, subjective success definition, observation window shorter than the payback period, forecast presented as observation, round total numbers, underpowered for the number of tests run, attribution drift, method misstated in coverage, and a dead or placeholder citation. The correctly sourced deskilling finding is stronger than the one that went viral. Endoscopists with a mean 28 years of practice saw unassisted adenoma detection fall from 28.4 percent to 22.4 percent after three months of routine AI exposure, adjusted odds ratio 0.69, 95 percent CI 0.53 to 0.89, p = 0.005. The Gartner release behind the 40 percent forecast also estimated that only about 130 of the thousands of agentic AI vendors are real. That is the more useful number and it got a fraction of the coverage.
Where does this data come from?
Every figure is reproduced from a named publisher: MIT NANDA, The GenAI Divide: State of AI in Business 2025, Kosmyna et al., Your Brain on ChatGPT, Budzyn et al., Lancet Gastroenterology and Hepatology 2025, Stanford Digital Economy Lab, Canaries in the Coal Mine. Sample, field date, and confidence are shown on each chart. Sources marked as vendor research are labelled on the page.
What could not be confirmed?
Forrester 88 percent and 22 percent. Present in no primary Forrester publication we could locate. Two aggregators report identical figures while disagreeing about which firm produced them. Do not chart, do not cite. Anthropic agent adoption at levels similar to OpenAI's. Anonymous sourcing only.

Cite this page

Permanent URL and suggested citation.

https://www.therevenueaireport.com/research/citation-decay

Kvarfordt, Jonathan. "The most-quoted AI numbers do not survive contact with their own methodology." The Revenue AI Report, Research Library. https://www.therevenueaireport.com/research/citation-decay

Figures on this page are reproduced from the publishers listed below. Cite the original publisher for the underlying data, and this page for the compilation and framing.

Sources

Every publisher used on this page.

If a metric, model term, or method on this page is unfamiliar, every one of them is defined in The AI and Revenue Dictionary. Sample size, field date, and confidence tags are explained there too.

Could not confirm

What we looked for and did not find.

Claims found during research and not charted

  • Forrester 88 percent and 22 percent. Present in no primary Forrester publication we could locate. Two aggregators report identical figures while disagreeing about which firm produced them. Do not chart, do not cite.
  • Anthropic agent adoption at levels similar to OpenAI's. Anonymous sourcing only.

How to cite this research

Written by Jonathan Kvarfordt, Founder and Principal Analyst, The Revenue AI Report. Published under CC BY 4.0.

APA

Kvarfordt, J. (2026). The most-quoted AI numbers do not survive contact with their own methodology. The Revenue AI Report. Retrieved from https://www.therevenueaireport.com/research/citation-decay

MLA

Kvarfordt, Jonathan. "The most-quoted AI numbers do not survive contact with their own methodology." The Revenue AI Report, 31 Aug. 2026, www.therevenueaireport.com/research/citation-decay.

BibTeX

@misc{kvarfordt2026citationdecay,
  author = {Kvarfordt, Jonathan},
  title = {The most-quoted AI numbers do not survive contact with their own methodology},
  year = {2026},
  publisher = {The Revenue AI Report},
  url = {https://www.therevenueaireport.com/research/citation-decay}
}

Share this research

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Next theme

The fear surge is mostly a coverage surge

Ten years of AI coverage measured directly. Volume rose 6.2x, tone fell 64 percent and never went negative, and fear coverage outran works coverage in exactly one year.

Subscribe

Get the weekly issue built on this data.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.