One condition separates agents that work from agents that do not
Vendors publish the first attempt. The fourth attempt is the deployment. Where a cheap automatic check fires before an irreversible action, agents hold. Where it does not, they collapse.
The short answer
What does the research show about One condition separates agents that work from agents that do not?
Evidence
- 34 percent of consumers would allow an agent to act only with per-action approval, 23 percent want suggestions only, and 21 percent want no agent action at all.
- Where there is no gate at all on an irreversible action, nine seconds of agent autonomy destroyed a production database and its backups.
- Sales outreach is the worst possible fit for the condition that predicts success. There is no cheap automatic check on whether an email should have been sent, and sending it is irreversible.
Supporting pages
- Text to Action, Not Text to Content: What Changes When Agents Execute analysis
- Dictionary plain-language definitions
pass^k is defined in the source paper as the fraction of k independent runs that succeed. Every task in the tau2-bench paper was run four times at temperature zero. The results are the most usable reliability numbers available, and not one 2026 vendor announcement fetched in this research publishes them.
The skeptic case, at full strength: OpenAI's Codex plus ChatGPT Work have about 10 million weekly users, described by the company as an estimate rather than an audited metric. Against roughly 900 million weekly ChatGPT users, agents are about 1.1 percent on a matched weekly basis. Josh Miller of The Browser Company called agents an invented frame, and the article that anchors this section flags his conflict itself: he sells an AI-powered browser.
The separating condition is not model capability. It is whether the task has a cheap, automatic correctness check that fires before any irreversible action, so a failed attempt costs one retry instead of one loss. Give the agent the correct plan and o4-mini's telecom pass^1 goes from 0.42 to 0.96. The capability was already there. The verification was what was missing.
What this page is
Multi-attempt agent reliability results at temperature zero, plus the single condition that separates agents that hold in production from agents that collapse.
The argument
Vendors publish the first attempt and deployments run on the fourth. Reliability is predicted by whether a cheap automatic check fires before an irreversible action, not by model capability.
How to read it
- pass^k is the fraction of k independent runs that succeed. Every task in the source paper was run four times at temperature zero.
- A single-attempt score tells you almost nothing about production behavior. The gap between pass^1 and pass^4 is the deployment risk.
- Usage comparisons here are company estimates rather than audited metrics, and are labeled that way.
Vendors publish the first attempt. The fourth attempt is the deployment.
tau2-bench, eight model and domain pairs, four independent runs per task at temperature zero.
0.02 by the fourth attempt on the mms_issue tasks
What this does not say
A higher score on the fourth attempt is not a better agent. It is the same agent given three more chances, which a production workflow may not allow.
- Publisher
- tau2-bench, arXiv 2506.07982, with the leaderboard figure from taubench.com
- Sample and method
- Four independent runs per task at temperature 0. Moving from autonomous operation to the collaborative default cost gpt-4.1 18 percent and o4-mini 25 percent of pass^1, and performance was close to zero beyond seven actions
- Field dates
- Not published by the source.
No 2026 vendor announcement fetched in this research publishes pass^k. Every published score is single-attempt.
Agents are about 1.1 percent of chatbot usage on a matched weekly basis.
Area-proportional squares, because two quantities four orders of magnitude apart cannot be read as angles.
- 900,000,000
- Weekly ChatGPT users, reported 27 February 2026
- 10,000,000
- Weekly Codex plus ChatGPT Work users, a company estimate
Both figures are weekly and both are company estimates. The 10 million is described by its own source as not an audited metric and excludes enterprise seats. The original published comparison used a monthly chatbot figure against a weekly agent figure. Corrected here to a matched weekly basis.
What this does not say
User counts are not usage. A monthly active user may have run one prompt.
- Publisher
- TechCrunch and TNW, from company statements
- Sample and method
- Gemini separately crossed 1 billion monthly users on 11 August 2026. Monthly and weekly bases are not mixed in this chart
- Field dates
- Not published by the source.
- Source
- https://techcrunch.com/
Both numbers are vendor self-reported. Neither is audited.
One condition predicts every agent outcome in the evidence: is there a free check before something irreversible happens?
Concept view with the numbers attached. There is no single dataset behind this comparison and the page says so.
Cheap automatic check before the irreversible action
Pull requests merged by an agent
34 to 67 percent in a year
Cognition, vendor reported
Touchless accounts payable
7 to about 65 percent
Genpact, humans validate exceptions
BrowseComp retrieval accuracy
51.5 to a claimed 92.2 percent
OpenAI, vendor reported
High-volume support resolution
50 to 90 percent
Vendor self-reported, none independently measured
No cheap check, or the artifact is the deliverable
Remote Labor Index automation rate
2.5 percent, $1,720 of $143,991 earned
arXiv 2510.26787
Vending-Bench 2 best agent
$10,936.76 against roughly $63,000 for a strong human
Andon Labs and Epoch AI
tau2-bench mms_issue, fourth attempt
0.02
arXiv 2506.07982
CRMArena-Pro, single turn to multi turn
about 58 to about 35 percent
arXiv 2505.18878
Consumers who have ordered via an assistant
6 percent
Dynata for Radial
Give the agent the correct plan and telecom pass^1 goes from 0.42 to 0.96 for o4-mini and 0.34 to 0.73 for gpt-4.1. The capability was already there.
What this does not say
Vendor-reported rows are not independent replications. They are labelled so the two kinds are not read together.
- Publisher
- Compiled by The Revenue AI Report
- Sample and method
- Every row carries its own publisher above. Vendor-reported rows are labelled as such
- Field dates
- Not published by the source.
- Source
- No primary URL reachable at research time.
The left column mixes vendor-reported deployment figures with peer-reviewed benchmarks. The condition is our framing; the numbers are as published.
Consumers are open to agents and have not used them.
Willingness sits at 58 percent. Every behavioural and trust measure sits under 20 percent.
Percent of respondents
- Start shopping with AI tools5%
Against 34 percent search and 32 percent marketplaces
- Have ordered via an AI assistant6%
- Trust AI to handle problems when something goes wrong15%
- Trust AI with payment data17%
- Trust AI to act in their financial interest18%
- Open to ordering via an AI assistant58%
58 percent open. 6 percent have done it. 60 percent would stop using an AI shopping agent after one mistake.
Top to bottom spread: 58 percent open against 6 percent who have done it
What this does not say
Stated consumer preference is not observed behavior. People report avoiding channels they still use.
- Publisher
- Dynata for Radial, and YouGov for ACI Worldwide · Vendor research
- Sample and method
- Two surveys of 1,000 US adults each, and YouGov n = 2,080 UK adults 18+, weighted
- Field dates
- December 2025, January 2026, and 19 to 22 June 2026
A competing 68 percent figure for consumers who have used a shopping agent is excluded: no sample size and no field dates. It is shown failing rather than omitted.
Willingness sits at 58 percent. Every behavioural and trust measure sits under 20 percent.
| Percent of respondents | Value (%) | Note |
|---|---|---|
| Start shopping with AI tools | 5 | Against 34 percent search and 32 percent marketplaces |
| Have ordered via an AI assistant | 6 | |
| Trust AI to handle problems when something goes wrong | 15 | |
| Trust AI with payment data | 17 | |
| Trust AI to act in their financial interest | 18 | |
| Open to ordering via an AI assistant | 58 | 58 percent open. 6 percent have done it. 60 percent would stop using an AI shopping agent after one mistake. |
Source: Dynata for Radial, and YouGov for ACI Worldwide. Two surveys of 1,000 US adults each, and YouGov n = 2,080 UK adults 18+, weighted Fielded December 2025, January 2026, and 19 to 22 June 2026. Confidence: Medium.
Vendor research. The publisher sells into the market it measured.
Caveat: A competing 68 percent figure for consumers who have used a shopping agent is excluded: no sample size and no field dates. It is shown failing rather than omitted.
What this does not say: Stated consumer preference is not observed behavior. People report avoiding channels they still use.
Also in the record
Figures that sit alongside these charts.
- 34 percent of consumers would allow an agent to act only with per-action approval, 23 percent want suggestions only, and 21 percent want no agent action at all.
- Where there is no gate at all on an irreversible action, nine seconds of agent autonomy destroyed a production database and its backups.
- Sales outreach is the worst possible fit for the condition that predicts success. There is no cheap automatic check on whether an email should have been sent, and sending it is irreversible.
- Ask any vendor you are evaluating for their pass^4. The refusal is the answer.
The brief
What is going on here, and why it matters.
The charts above are the evidence. This is the read: what the data describes, the mechanism behind it, where the argument could be wrong, and what a revenue team does about it.
The first attempt is not the product
Every task in the tau2-bench paper was run four times at temperature zero, and by the fourth attempt agents without a cheap automatic check before an irreversible action collapse. Those are the most usable reliability numbers publicly available, and not one 2026 vendor announcement fetched during this research publishes them.
That omission is the whole procurement lesson. Ask any vendor you are evaluating for their pass^4. The refusal is the answer, and it is a faster diligence signal than any demo.
The separating condition
It is not model capability. It is whether the task has a cheap, automatic correctness check that fires before any irreversible action, so a failed attempt costs one retry instead of one loss. Give the agent the correct plan and o4-mini's telecom pass^1 goes from 0.42 to 0.96. The capability was already there. The verification was what was missing.
The negative case makes the same point at full cost. Where there was no gate at all on an irreversible action, nine seconds of agent autonomy destroyed a production database and its backups. Autonomy without a pre-commit check is not a capability level, it is an unbounded loss.
Why sales outreach is the worst fit
Sales outreach fails the condition on both halves. There is no cheap automatic check on whether an email should have been sent, and sending it is irreversible. That is the exact profile the evidence says collapses, and it is also the single most-funded agent category in GTM.
The workable version is to move the gate earlier: the agent drafts and a human commits, or the agent acts only inside a reversible surface such as an internal record, a draft queue, or a research summary. Those designs keep the throughput and remove the unbounded loss.
Scale check
OpenAI's Codex plus ChatGPT Work have about 10 million weekly users, described by the company as an estimate rather than an audited metric. Against roughly 900 million weekly ChatGPT users, agents are about 1.1 percent on a matched weekly basis. The agent market is early by usage, whatever the roadmap says.
Consumer posture matches. Thirty-four percent would allow an agent to act only with per-action approval, 23 percent want suggestions only, and 21 percent want no agent action at all. Openness is high, delegation is not.
What to do with it
The move, by seat.
- CRO
- Ask every agent vendor for pass^4 at temperature zero. No answer is an answer.
- RevOps
- Put a cheap automatic check before every irreversible action, or make the action reversible. That one design choice predicts the outcome.
- Sales leaders
- Keep outbound send behind a human commit. Outreach fails both halves of the condition that predicts agent success.
Questions this page answers
What the data says, in plain language.
- What does the research show about One condition separates agents that work from agents that do not?
- Vendors publish the first attempt. The fourth attempt is the deployment. Where a cheap automatic check fires before an irreversible action, agents hold. Where it does not, they collapse. pass^k is defined in the source paper as the fraction of k independent runs that succeed. Every task in the tau2-bench paper was run four times at temperature zero. The results are the most usable reliability numbers available, and not one 2026 vendor announcement fetched in this research publishes them.
- What does the figure "Vendors publish the first attempt. The fourth attempt is the deployment" show?
- tau2-bench, eight model and domain pairs, four independent runs per task at temperature zero. Source: tau2-bench, arXiv 2506.07982, with the leaderboard figure from taubench.com. Four independent runs per task at temperature 0. Moving from autonomous operation to the collaborative default cost gpt-4.1 18 percent and o4-mini 25 percent of pass^1, and performance was close to zero beyond seven actions Confidence: High. Caveat: No 2026 vendor announcement fetched in this research publishes pass^k. Every published score is single-attempt.
- What does the figure "Agents are about 1.1 percent of chatbot usage on a matched weekly basis" show?
- Area-proportional squares, because two quantities four orders of magnitude apart cannot be read as angles. Source: TechCrunch and TNW, from company statements. Gemini separately crossed 1 billion monthly users on 11 August 2026. Monthly and weekly bases are not mixed in this chart Confidence: Medium. Caveat: Both numbers are vendor self-reported. Neither is audited.
- What does the figure "One condition predicts every agent outcome in the evidence: is there a free check before something irreversible happens?" show?
- Concept view with the numbers attached. There is no single dataset behind this comparison and the page says so. Source: Compiled by The Revenue AI Report. Every row carries its own publisher above. Vendor-reported rows are labelled as such Confidence: Medium. Caveat: The left column mixes vendor-reported deployment figures with peer-reviewed benchmarks. The condition is our framing; the numbers are as published.
- What does the figure "Consumers are open to agents and have not used them" show?
- Willingness sits at 58 percent. Every behavioural and trust measure sits under 20 percent. Source: Dynata for Radial, and YouGov for ACI Worldwide. Two surveys of 1,000 US adults each, and YouGov n = 2,080 UK adults 18+, weighted Fielded December 2025, January 2026, and 19 to 22 June 2026. Confidence: Medium. Caveat: A competing 68 percent figure for consumers who have used a shopping agent is excluded: no sample size and no field dates. It is shown failing rather than omitted. This figure comes from vendor research, so read it as a vendor claim.
- What else sits alongside these figures?
- 34 percent of consumers would allow an agent to act only with per-action approval, 23 percent want suggestions only, and 21 percent want no agent action at all. Where there is no gate at all on an irreversible action, nine seconds of agent autonomy destroyed a production database and its backups. Sales outreach is the worst possible fit for the condition that predicts success. There is no cheap automatic check on whether an email should have been sent, and sending it is irreversible. Ask any vendor you are evaluating for their pass^4. The refusal is the answer.
- Where does this data come from?
- Every figure is reproduced from a named publisher: arXiv 2506.07982, tau2-bench, arXiv 2406.12045, the original tau-bench, arXiv 2510.26787, Remote Labor Index, arXiv 2505.18878, CRMArena-Pro, arXiv 2307.13854, WebArena, WIRED, Why Normal People Aren't Using AI Agents. Sample, field date, and confidence are shown on each chart. Sources marked as vendor research are labelled on the page.
- What could not be confirmed?
- Anthropic agent adoption at levels similar to OpenAI's. Anonymous sourcing only, not charted. 68 percent of consumers having already used an AI shopping agent. No sample size, no field dates.
Cite this page
Permanent URL and suggested citation.
https://www.therevenueaireport.com/research/agent-reliability
Kvarfordt, Jonathan. "One condition separates agents that work from agents that do not." The Revenue AI Report, Research Library. https://www.therevenueaireport.com/research/agent-reliability
Figures on this page are reproduced from the publishers listed below. Cite the original publisher for the underlying data, and this page for the compilation and framing.
Sources
Every publisher used on this page.
If a metric, model term, or method on this page is unfamiliar, every one of them is defined in The AI and Revenue Dictionary. Sample size, field date, and confidence tags are explained there too.
arXiv 2506.07982, tau2-bench
Four independent runs per task at temperature 0, per-model and per-domain pass^k degradation
https://arxiv.org/abs/2506.07982arXiv 2406.12045, the original tau-bench
Defines pass^k as the fraction of k independent runs that succeed
https://arxiv.org/abs/2406.12045arXiv 2510.26787, Remote Labor Index
Best automation rate 2.5 percent, $1,720 earned out of $143,991 available
https://arxiv.org/abs/2510.26787arXiv 2505.18878, CRMArena-Pro
About 58 percent single-turn falling to about 35 percent multi-turn, with near-zero confidentiality awareness under a standard prompt
https://arxiv.org/abs/2505.18878arXiv 2307.13854, WebArena
Best GPT-4-based agent 14.41 percent against humans at 78.24 percent
https://arxiv.org/abs/2307.13854WIRED, Why Normal People Aren't Using AI Agents
Maxwell Zeff, 6 August 2026. Anchor article for the skeptic case, including its own conflict disclosure
https://www.wired.com/
Could not confirm
What we looked for and did not find.
Claims found during research and not charted
- Anthropic agent adoption at levels similar to OpenAI's. Anonymous sourcing only, not charted.
- 68 percent of consumers having already used an AI shopping agent. No sample size, no field dates.
Read the analysis
Issues built on this theme.
Research on this site is the evidence layer. These essays take the numbers above and apply them to real decisions, so you can see how the data reads in practice.
- Text to Action, Not Text to Content: What Changes When Agents Execute
The shift is not better output. It is AI doing the next step. What a revenue leader has to own when the tool stops drafting and starts executing.
How to cite this research
Written by Jonathan Kvarfordt, Founder and Principal Analyst, The Revenue AI Report. Published under CC BY 4.0.
APA
Kvarfordt, J. (2026). One condition separates agents that work from agents that do not. The Revenue AI Report. Retrieved from https://www.therevenueaireport.com/research/agent-reliability
MLA
Kvarfordt, Jonathan. "One condition separates agents that work from agents that do not." The Revenue AI Report, 31 Aug. 2026, www.therevenueaireport.com/research/agent-reliability.
BibTeX
@misc{kvarfordt2026agentreliability,
author = {Kvarfordt, Jonathan},
title = {One condition separates agents that work from agents that do not},
year = {2026},
publisher = {The Revenue AI Report},
url = {https://www.therevenueaireport.com/research/agent-reliability}
}Next theme
What the peer-reviewed literature says, rankedTen studies ranked on identification strength, sample size, outcome relevance, and peer-review status together. The strongest findings contradict the vendor narrative in ten specific places.
Subscribe
