The Task Fallacy

Why Anthropic's extreme scenario assumes away the actual job. Eight named reversals, one enterprise return record, and one model with the reinstatement effect turned off.

The Task Fallacy: a job is not just a bag of tasks. A visual summary contrasting the Anthropic task-automation assumption with the Revenue AI Report view that jobs retain trust, judgment, and accountability.
Source: The Revenue AI Report. Cite: https://www.therevenueaireport.com/research/task-fallacy

Share

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

The short answer

What does the research show about The Task Fallacy?

The task fallacy is the mistake of treating a job as a bundle of tasks that can each be automated, then counting the tasks as the job. Research from the Anthropic Institute, MIT, BCG, and our own Reversal Ledger shows that automating the artifact does not remove the accountability. When accountability returns to a person, the company pays twice: once for the tool, and once for the human who has to clean it up.

Evidence

  • Buy against a job, not a task. Write the artifact the tool produces, then write who signs for it. If the signer is still a person, the headcount case is a cost case, not a replacement case.
  • Price the accountability. Every reversal in the ledger paid twice, once for the automation and once for the human who came back. Put the rehire cost in the business case before the pilot starts.
  • Set a kill criterion at purchase. Air Canada, DPD, and Cursor all discovered the criterion after the incident. Name the failure that ends the deployment and the person who calls it.

Supporting pages

On 9 September 2026 the Anthropic Institute published Economic Scenarios Working Paper No. 2026-02. The paper runs a model forward to 2030 under three published parameter settings. The extreme setting produced the headline numbers: 17.9% cognitive-worker unemployment, wages 11.5% below the no-AI path, and labor share of income falling from 60% to 45.2%. Those numbers are outputs of a model, not measurements of the economy. The publisher summary explains the three scenarios and states that they are not predictions.

The model reaches those outputs by turning three economic mechanisms to their extremes. The reinstatement ratio is set to 0.00, meaning no new work is created as old tasks disappear. The automation share is set to 0.90, meaning nine in ten AI-performed tasks happen with no human involved. The search discount is set to 0.04, meaning displaced workers barely move between occupations. Every prior wave of automation was absorbed by exactly those three mechanisms, so the extreme scenario is a useful stress test, not a base case.

The gap between a task and a job is what this page tests. A task produces an artifact. A job carries trust, judgment, and accountability for that artifact. The Reversal Ledger shows the same pattern across eight named cases from 2017 to 2025: the artifact was automated, the accountability came back to a person, and the company paid twice.

Enterprise return data matches the ledger, not the extreme scenario. MIT's NANDA initiative found 95% of generative AI pilots delivered zero measurable P&L return. BCG found 74% of companies struggle to scale AI value beyond proof of concept, and only 5% qualify as AI-future built. The charts below explain what the Anthropic model assumes, how real returns look, and which seats are most exposed to the task fallacy.

Eight named reversals. Every one automated what the task frame said was automatable.

Publicly documented reversals of AI task automation, February 2017 to May 2025, in lanes by the kind of event. Klarna is in copper because it is the one case that put humans back at scale.

Shut off or rolled back

  • IBM Watson at MD AndersonFeb 2017

    Oncology advisory project halted after audit, system never used on patients in the pilot scope

  • DPDJan 2024

    Support chatbot swore at a customer and criticised its own employer, AI element disabled

  • Air CanadaFeb 2024

    Tribunal held the airline liable for the chatbot's bereavement-fare answer, the bot came down

  • McDonald's with IBMJul 2024

    Drive-thru voice ordering ended across more than 100 restaurants after order-accuracy failures

  • CursorApr 2025

    Support bot invented a login policy that did not exist, cancellations followed, human review restored

Replaced humans, then rehired

  • KlarnaMay 2025

    CEO said cost had become a too predominant factor, quality fell, human agents brought back

Public walk-back of the message

  • DuolingoApr 2025

    AI-first announcement walked back publicly after user and staff reaction

Program wound down at a loss

  • Zillow OffersNov 2021

    Algorithmic home buying wound down, $421.6M segment loss before tax and roughly a quarter of staff cut

What this does not say

This is a collected set of reported cases, not a sample. It cannot tell you what share of all AI deployments get reversed.

Publisher
Bloomberg, Fortune, BBC, Ars Technica, The Register, Forbes, Wired, Restaurant Business, ITV
Sample and method
Eight publicly documented cases collected by editorial. Not a probability sample.
Field dates
Events dated 19 February 2017 to 9 May 2025
Medium confidence

Enterprise AI is not producing the value the extreme scenario requires.

Three surveys, 2024 to 2025, of what corporate AI actually returns.

Share of companies reporting the outcome

  • MIT NANDA: generative AI pilots delivering zero measurable P&L return95%
  • BCG: companies struggling to scale AI value beyond proof of concept74%
  • BCG: companies described as AI-future built5%

What this does not say

A pilot delivering no P&L return is not the same as a failed pilot. Some pilots are learning investments that pay out later.

Publisher
MIT NANDA initiative, Boston Consulting Group
Sample and method
MIT: 150 leader interviews, 350 employee surveys, analysis of 300 public deployments. BCG figures: global executive surveys.
Field dates
MIT published August 2025. BCG scale-value report published October 2024. BCG AI-future built report published 2025. Field windows not fully specified by the publishers.
Medium confidence
Enterprise AI is not producing the value the extreme scenario requires.

Three surveys, 2024 to 2025, of what corporate AI actually returns.

Enterprise AI is not producing the value the extreme scenario requires.
Share of companies reporting the outcomeValue (%)Note
MIT NANDA: generative AI pilots delivering zero measurable P&L return95
BCG: companies struggling to scale AI value beyond proof of concept74
BCG: companies described as AI-future built5

Source: MIT NANDA initiative, Boston Consulting Group. MIT: 150 leader interviews, 350 employee surveys, analysis of 300 public deployments. BCG figures: global executive surveys. Fielded MIT published August 2025. BCG scale-value report published October 2024. BCG AI-future built report published 2025. Field windows not fully specified by the publishers.. Confidence: Medium.

What this does not say: A pilot delivering no P&L return is not the same as a failed pilot. Some pilots are learning investments that pay out later.

The extreme scenario is a dial setting, not a forecast.

Parameter values from Anthropic Working Paper 2026-02, Table 2, evaluated at the start of 2030. The substantial scenario sits between the two columns.

Extreme scenario settings

  • Reinstatement ratio, new labor tasks created per automated task

    0.00

    no new work created

  • Automation share, fraction of AI work done without a human

    0.90

    nine in ten tasks unsupervised

  • Search discount, effectiveness of cross-occupation job search

    0.04

    workers barely move

  • Posting speed, monthly absorption by other occupations

    0.50

  • Affected task mass, share of all tasks AI can reach by 2030

    0.50

Modest scenario settings

  • Reinstatement ratio

    0.50

    new work created

  • Automation share

    0.50

    half the tasks keep a human

  • Search discount

    0.17

    workers move between occupations

  • Posting speed

    0.10

  • Affected task mass

    not published for this scenario

Extreme is the setting where the reinstatement effect is turned off and 90% of AI-performed work happens with no human. Those are choices, not forecasts.

What this does not say

The paper does not attach probabilities to these scenarios and explicitly states they are not predictions.

Publisher
The Anthropic Institute
Sample and method
Working Paper No. 2026-02, Table 2, parameter values for the published scenarios.
Field dates
Paper published 9 September 2026. Base period 2024, anchor mid-2026, evaluation start of 2030.
High confidence

The more the job runs on trust, the wider the gap between task automation and job replacement.

Editorial estimates by role of the share of the wage that pays for trust, judgment, and accountability rather than artifact production.

Share of the role that is trust, judgment, and accountability

  • SDR and outbound25%

    75% artifact production

  • Marketing content30%

    70% artifact production

  • RevOps analyst40%

    60% artifact production

  • Account executive60%

    40% artifact production

  • Customer success manager70%

    30% artifact production

  • Enablement and coaching75%

    25% artifact production

  • Finance close and audit sign-off80%

    20% artifact production

  • Founder-led sale90%

    10% artifact production

What this does not say

These are editorial estimates for illustration, not measured splits from a survey. Actual proportions vary by company, deal size, and buyer segment.

Publisher
The Revenue AI Report
Sample and method
Editorial estimate based on the reversal ledger and the role composition of enterprise revenue teams.
Field dates
Compiled September 2026
Low confidence
The more the job runs on trust, the wider the gap between task automation and job replacement.

Editorial estimates by role of the share of the wage that pays for trust, judgment, and accountability rather than artifact production.

The more the job runs on trust, the wider the gap between task automation and job replacement.
Share of the role that is trust, judgment, and accountabilityValue (%)Note
SDR and outbound2575% artifact production
Marketing content3070% artifact production
RevOps analyst4060% artifact production
Account executive6040% artifact production
Customer success manager7030% artifact production
Enablement and coaching7525% artifact production
Finance close and audit sign-off8020% artifact production
Founder-led sale9010% artifact production

Source: The Revenue AI Report. Editorial estimate based on the reversal ledger and the role composition of enterprise revenue teams. Fielded Compiled September 2026. Confidence: Low.

What this does not say: These are editorial estimates for illustration, not measured splits from a survey. Actual proportions vary by company, deal size, and buyer segment.

Three scenarios. One shifted assumption at a time.

Same Anthropic model, three published parameter sets, five outputs at 2030. Read each panel left to right across the scenarios.

2030 GDP versus the no-AI path · percent

32.40

Economy-wide unemployment · percent

11.90

Cognitive-worker unemployment · percent, modest not published

17.90

Cognitive-worker wages versus no-AI · percent, zero line is the no-AI path

no-AI path0.4-11.5

Labor share of income · percent, no-AI baseline 60

59.40
ModestSubstantialExtreme

The extreme scenario produces its results only when the reinstatement ratio is 0.00, the automation share is 0.90, and the search discount is 0.04. Real-economy evidence for those settings is absent.

What this does not say

These are outputs of the Anthropic model at the parameter values the authors chose for each scenario. They are not empirical measurements of 2030.

Publisher
The Anthropic Institute
Sample and method
Working Paper No. 2026-02, published scenario outputs for 2030.
Field dates
Paper published 9 September 2026, scenario values evaluated at the start of 2030
High confidence

Also in the record

Figures that sit alongside these charts.

  • Buy against a job, not a task. Write the artifact the tool produces, then write who signs for it. If the signer is still a person, the headcount case is a cost case, not a replacement case.
  • Price the accountability. Every reversal in the ledger paid twice, once for the automation and once for the human who came back. Put the rehire cost in the business case before the pilot starts.
  • Set a kill criterion at purchase. Air Canada, DPD, and Cursor all discovered the criterion after the incident. Name the failure that ends the deployment and the person who calls it.
  • Treat the extreme scenario as a stress test. Ask what your plan looks like at a reinstatement ratio of zero, then ask what evidence you have that your own function is anywhere near that setting.
  • Measure the trust share of each seat before you automate it. The higher that share, the smaller the headcount effect of automating the artifact, and the more the tool has to justify itself on cycle time rather than on people.

Questions this page answers

What the data says, in plain language.

What does the research show about The Task Fallacy?
Why Anthropic's extreme scenario assumes away the actual job. Eight named reversals, one enterprise return record, and one model with the reinstatement effect turned off. On 9 September 2026 the Anthropic Institute published Economic Scenarios Working Paper No. 2026-02. The paper runs a model forward to 2030 under three published parameter settings. The extreme setting produced the headline numbers: 17.9% cognitive-worker unemployment, wages 11.5% below the no-AI path, and labor share of income falling from 60% to 45.2%. Those numbers are outputs of a model, not measurements of the economy. The publisher summary explains the three scenarios and states that they are not predictions.
What does the figure "Eight named reversals. Every one automated what the task frame said was automatable" show?
Publicly documented reversals of AI task automation, February 2017 to May 2025, in lanes by the kind of event. Klarna is in copper because it is the one case that put humans back at scale. Source: Bloomberg, Fortune, BBC, Ars Technica, The Register, Forbes, Wired, Restaurant Business, ITV. Eight publicly documented cases collected by editorial. Not a probability sample. Fielded Events dated 19 February 2017 to 9 May 2025. Confidence: Medium.
What does the figure "Enterprise AI is not producing the value the extreme scenario requires" show?
Three surveys, 2024 to 2025, of what corporate AI actually returns. Source: MIT NANDA initiative, Boston Consulting Group. MIT: 150 leader interviews, 350 employee surveys, analysis of 300 public deployments. BCG figures: global executive surveys. Fielded MIT published August 2025. BCG scale-value report published October 2024. BCG AI-future built report published 2025. Field windows not fully specified by the publishers.. Confidence: Medium.
What does the figure "The extreme scenario is a dial setting, not a forecast" show?
Parameter values from Anthropic Working Paper 2026-02, Table 2, evaluated at the start of 2030. The substantial scenario sits between the two columns. Source: The Anthropic Institute. Working Paper No. 2026-02, Table 2, parameter values for the published scenarios. Fielded Paper published 9 September 2026. Base period 2024, anchor mid-2026, evaluation start of 2030.. Confidence: High.
What does the figure "The more the job runs on trust, the wider the gap between task automation and job replacement" show?
Editorial estimates by role of the share of the wage that pays for trust, judgment, and accountability rather than artifact production. Source: The Revenue AI Report. Editorial estimate based on the reversal ledger and the role composition of enterprise revenue teams. Fielded Compiled September 2026. Confidence: Low.
What else sits alongside these figures?
Buy against a job, not a task. Write the artifact the tool produces, then write who signs for it. If the signer is still a person, the headcount case is a cost case, not a replacement case. Price the accountability. Every reversal in the ledger paid twice, once for the automation and once for the human who came back. Put the rehire cost in the business case before the pilot starts. Set a kill criterion at purchase. Air Canada, DPD, and Cursor all discovered the criterion after the incident. Name the failure that ends the deployment and the person who calls it. Treat the extreme scenario as a stress test. Ask what your plan looks like at a reinstatement ratio of zero, then ask what evidence you have that your own function is anywhere near that setting.
Where does this data come from?
Every figure is reproduced from a named publisher: Anthropic Institute, Economic Scenarios Working Paper No. 2026-02, Anthropic Institute, economic scenarios overview, MIT NANDA initiative, coverage by Fortune, Boston Consulting Group, AI adoption in 2024, Boston Consulting Group, are you generating value from AI, The Revenue AI Report, Reversal Ledger. Sample, field date, and confidence are shown on each chart. Sources marked as vendor research are labelled on the page.
What could not be confirmed?
A published probability for any of the three Anthropic scenarios. The paper states the scenarios are not predictions and attaches no likelihoods. A measured reinstatement ratio for enterprise revenue functions. No publisher has produced one, so the trust-share chart on this page is an editorial estimate and is labeled as such.

Cite this page

Permanent URL and suggested citation.

https://www.therevenueaireport.com/research/task-fallacy

Kvarfordt, Jonathan. "The Task Fallacy." The Revenue AI Report, Research Library. https://www.therevenueaireport.com/research/task-fallacy

Figures on this page are reproduced from the publishers listed below. Cite the original publisher for the underlying data, and this page for the compilation and framing.

Sources

Every publisher used on this page.

If a metric, model term, or method on this page is unfamiliar, every one of them is defined in The AI and Revenue Dictionary. Sample size, field date, and confidence tags are explained there too.

Could not confirm

What we looked for and did not find.

Claims found during research and not charted

  • A published probability for any of the three Anthropic scenarios. The paper states the scenarios are not predictions and attaches no likelihoods.
  • A measured reinstatement ratio for enterprise revenue functions. No publisher has produced one, so the trust-share chart on this page is an editorial estimate and is labeled as such.

How to cite this research

Written by Jonathan Kvarfordt, Founder and Principal Analyst, The Revenue AI Report. Published under CC BY 4.0.

APA

Kvarfordt, J. (2026). The Task Fallacy. The Revenue AI Report. Retrieved from https://www.therevenueaireport.com/research/task-fallacy

MLA

Kvarfordt, Jonathan. "The Task Fallacy." The Revenue AI Report, 31 Aug. 2026, www.therevenueaireport.com/research/task-fallacy.

BibTeX

@misc{kvarfordt2026taskfallacy,
  author = {Kvarfordt, Jonathan},
  title = {The Task Fallacy},
  year = {2026},
  publisher = {The Revenue AI Report},
  url = {https://www.therevenueaireport.com/research/task-fallacy}
}

Share this research

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Next theme

The Proof Gap has a measured size

Every GTM function adopted AI faster than it produced revenue. One dataset measures both sides in the same sample.

Subscribe

Get the weekly issue built on this data.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.