Reality Check

What Truth Tests Should We Run Before Buying Claudeforce, Agentforce, or an AI SDR?

AI on a dirty CRM is a prettier lie. Five truth tests a Director+ can run before Claudeforce, Agentforce, or AI SDR spend. Definitions and freshness first. Then the pilot.

Jonathan Kvarfordt · Published August 27, 2026 · 10 min read

Why trust this analysis?

The short answer

Will Claudeforce or Agentforce fix a lying pipeline?

No. Both read your CRM. Claudeforce skills like deal health and pipeline review summarize the record you already have, so drift in stages, next steps, and amounts comes back as a confident summary of the wrong thing. Fix definitions and freshness first, then pilot.

Evidence

  • What separates the deployments that work The largest gap between AI leaders and everyone else is not technology. It is having decided what to build.
  • What is a pipeline truth test? A query plus a pass threshold that tells you whether the CRM matches reality. The five in this issue are stage definition, next-step freshness, amount integrity, closed-lost coding, and write-back ownership. Each one can be run on the last two closed quarters plus open pipeline.

Supporting pages

Last reviewed

AI on a dirty CRM is a prettier lie. Marc Benioff put it plainly on the Salesforce Q2 FY27 earnings call: if the data isn't right, the AI isn't right. Every layer you add on top of a pipeline that already misreports reality makes the misreporting faster, more confident, and harder to argue with.

This is not a tooling complaint. It is a sequencing decision. Before you sign for Claudeforce, expand Agentforce, or run an AI SDR pilot, you need to know whether the pipeline those systems will read is telling the truth. The GTM AI Podcast has covered the pipeline-lying problem from the practitioner side as well. This issue is the test protocol.

The argument

How this reality check breaks down

A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.

Contents diagram for What Truth Tests Should We Run Before Buying Claudeforce, Agentforce, or an AI SDR?, listing the sections: What lying actually means, Why this shows up as an AI outcome problem, The five truth tests, If three or more fail, If they pass.

What lying actually means

Lying is not fraud. It is drift. Six specific places where the record and the reality separate:

  • Stages. The stage name exists, the exit criteria do not, so two reps in the same segment call the same deal two different things.
  • Next steps. A field that is blank, stale, or filled with following up is not a next step. It is a placeholder that an AI will read as activity.
  • Amounts. List price, expected ARR, and total contract value get entered into whichever field is closest, then rolled up as if they mean the same thing.
  • Closed-lost. One picklist value absorbs everything: no decision, budget freeze, lost to a named competitor, lost to internal build, disqualified late.
  • Contact roles. Nobody records who the economic buyer was, so no model can learn what a winnable buying group looks like.
  • Forecast category. Commit, best case, and pipeline get set by rep instinct rather than by a rule, so the roll-up is a mood, not a measurement.

The test

Five questions a pipeline has to survive before AI can read it

Run in order. A failure at any rung invalidates the rungs above. Schematic, not a dataset. Source-cited charts live in the research library.

Five questions a pipeline has to survive before AI can read it. Diagram showing Is the stage defined?, Is the stage earned?, Is the date real?, Is the buyer engaged?, Would you bet on it?.

Each of those is survivable when a human sales manager is reading the record and applying judgment. None of them survive automation. An agent does not know the difference between a deal that is real and a deal that is well-formatted.

Why this shows up as an AI outcome problem

Gong's conversation data shows 81 percent of deals stall at least once, and stalls are where the record and the reality separate hardest. Gartner's finding on the other side: teams with clean CRM data see roughly 2x better AI forecast accuracy than teams without it. Both numbers are in our issue on GTM loop engineering, which is the argument for putting AI inside a closed loop rather than beside one.

The architectural version of the same point is in The Headless GTM Stack: when the UI becomes optional and sellers work from Claude, Slack, or an agent surface, the system of record stops being something a human eyeballs daily. Nobody is reading the row anymore. So the row has to be right by construction.

The five truth tests

Each test is a query plus a threshold. Run them on the last two closed quarters plus current open pipeline. A test fails if it misses the threshold. Two failures is a warning. Three or more is a delay.

Test 1: Stage definition

Pull your stage list and ask three managers independently to write the exit criteria for each stage from memory. Compare. Pass condition: the three answers match on every stage and match the documented definition. If there is no documented definition, the test has already failed. Then check the data: what percentage of deals skipped a stage entirely, and what percentage moved backward? Stage-skip above 20 percent means the stages are labels, not gates.

Test 2: Next-step freshness

For all open deals in the current quarter, measure the age of the next-step field and the presence of a real date. Pass condition: 85 percent of open deals have a next step updated in the last 14 days with a specific date and a named person. Count check in and follow up as blank. This test is the single best proxy for whether an AI deal-health skill will produce anything trustworthy, because deal health is mostly a function of recency.

Test 3: Amount integrity

Take 30 closed-won deals at random and reconcile the CRM amount against the signed order form or the billing system. Pass condition: 90 percent match within 5 percent. Then check the open pipeline for the tells: round numbers repeated across many deals, amounts unchanged from creation to close, and amounts entered before any pricing conversation happened. An AI pipeline review that sums a fictional number produces a fictional forecast with a confidence score attached.

Test 4: Closed-lost coding

Count the distinct closed-lost reasons actually used in the last two quarters, and the share of losses sitting in the top value. Pass condition: at least five reasons in genuine use, with no single value above 40 percent, and a competitor named where a competitor won. If no decision or other absorbs the majority, you have no loss signal. Every model you buy will learn what winning looks like and nothing about what losing looks like, which is the half that changes behavior.

Test 5: Write-back ownership

For every system that will write to the CRM after the pilot, name the owner, the field, the trigger, and the conflict rule in one document. Pass condition: no field has two writers without a documented precedence rule, and every automated write is attributable to a source. This is the test most teams skip and the one that turns a working pilot into an unrecoverable data problem two quarters later. Agent write-back without ownership is how a clean CRM becomes a dirty one at machine speed.

If three or more fail

Delay the purchase. Not forever, and not as a moral position: you are delaying because the pilot cannot produce a readable result. If the underlying data is drifting, a pilot that looks good and a pilot that looks bad tell you the same amount, which is nothing. You will spend a quarter arguing about the tool when the argument is about the fields.

Do this instead, in a 30-day window. Week one: write and publish stage exit criteria, and rebuild the closed-lost picklist. Week two: enforce next-step required-on-stage-advance and backfill current quarter. Week three: reconcile amounts and fix the field-of-record. Week four: write the write-back ownership document and get the RevOps lead and the systems owner to sign it. Then rerun the five tests and pilot on the other side.

If they pass

Run the pilot, and scope it to the tests you passed. If stage definition and next-step freshness are clean, deal health and pipeline review skills have something real to read. If amount integrity is clean, forecast assistance is worth measuring. If closed-lost coding is thin, do not buy a tool whose pitch is loss analysis, no matter how good the demo looks.

Then keep the tests. Rerun all five at the end of each quarter and treat a drop below threshold as a production incident, not a hygiene chore. The pipeline does not stay honest on its own, and neither does an agent reading it.

Take it to the room

The short list this issue leaves you with

Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.

Checklist diagram summarising What Truth Tests Should We Run Before Buying Claudeforce, Agentforce, or an AI SDR?: Stages; Next steps; Amounts; Closed-lost; Contact roles.

Frequently asked questions

Will Claudeforce or Agentforce fix a lying pipeline?
No. Both read your CRM. Claudeforce skills like deal health and pipeline review summarize the record you already have, so drift in stages, next steps, and amounts comes back as a confident summary of the wrong thing. Fix definitions and freshness first, then pilot.
What is a pipeline truth test?
A query plus a pass threshold that tells you whether the CRM matches reality. The five in this issue are stage definition, next-step freshness, amount integrity, closed-lost coding, and write-back ownership. Each one can be run on the last two closed quarters plus open pipeline.
When should we delay an AI purchase?
When three or more of the five tests fail. Not as a policy stance, but because a pilot running on drifting data produces a result nobody can read. Spend 30 days on definitions, freshness, amounts, and write-back ownership, rerun the tests, then pilot.
Does a clean CRM really double AI forecast accuracy?
Gartner's finding is roughly 2x better AI forecast accuracy for teams with clean CRM data versus those without. Gong's conversation data adds that 81 percent of deals stall at least once, which is where the record and the reality tend to separate.
Does a headless GTM stack change these tests?
It raises the stakes. When sellers work from Claude, Slack, or an agent surface, nobody eyeballs the CRM row daily, so errors stop being caught by human review. The record has to be correct by construction, which makes write-back ownership the most important of the five tests.
Who should own the truth tests?
RevOps owns the queries and thresholds, sales leadership owns the stage and forecast-category definitions, and the systems owner signs the write-back document. Run all five at the end of each quarter and treat a threshold miss as an incident rather than a cleanup task.

Share this issue

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Subscribe

Get the next Reality Check before you sign the order form.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.

Keep reading

Reality Check

Agentforce Pricing Explained: Credits, Licenses, and Total Cost

Agentforce is not one price. It is a stack of editions, entitlements, meters, and platform costs that only resolve into a number once you know which product you are buying. Here is how to work out which one you are looking at, and which question to ask next.