Why Your AI Pilot Worked and Your Rollout Did Not
The pilot ran on your best rep, your cleanest data, and your most motivated manager. Production runs on none of those. Here is what breaks between the two, and how to design a pilot that survives contact with the org.
Jonathan Kvarfordt · Published June 16, 2026 · 11 min read
The short answer
Why do AI pilots succeed but rollouts fail?
Evidence
- What separates the deployments that work The largest gap between AI leaders and everyone else is not technology. It is having decided what to build.
- How do you design an AI pilot that predicts production results? Run two cohorts, volunteers and assigned participants, and report the assigned cohort's numbers. Use unfiltered live data, capture a baseline before day one, name an owner with an ongoing budget line, and write kill criteria before launch.
Supporting pages
- What separates the deployments that work the data behind this piece
- The Proof Gap definition
- The Single-Player AI Problem definition
Last reviewed
The pilot slide always looks the same. Twelve reps, six weeks, a double-digit lift in something. The room nods. Budget moves. Six months later the same capability is live for four hundred people and nobody can find the lift.
This is not a tooling failure. It is a sampling failure. The pilot measured a system that no longer exists once you scale it, because almost every variable that made the pilot work was removed on the way to production.
The argument
How this post-mortem breaks down
A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.
Contents diagram for Why Your AI Pilot Worked and Your Rollout Did Not, listing the sections: Substitution one: you swapped your best peopl…, Substitution two: you swapped curated data fo…, Substitution three: you swapped attention for…, Substitution four: you swapped a narrow use c…, What a production-grade pilot looks like, The question to ask in the next steering meet….If you are about to green-light a rollout on the strength of a pilot, read the four substitutions below first. They are the difference between a number you can defend to a board and a number you will quietly stop reporting.
Substitution one: you swapped your best people for your average people
Pilots are staffed with volunteers. Volunteers are, by definition, the segment of your org most motivated to make a new tool work. They tolerate rough edges, they write better prompts, they report bugs instead of abandoning the workflow.
The gap
The five gates between a good pilot and a production workflow
Most stalls happen at the gate teams skipped. Schematic, not a dataset. Source-cited charts live in the research library.
The five gates between a good pilot and a production workflow. Diagram showing Pilot wins, Owner named, Data fixed, Process rewritten, In production.Production is staffed with everyone. The rep who is behind on quota, the manager who did not want the change, the enterprise AE who already has a working process and no incentive to risk it in Q4. The same tool in those hands produces a different distribution of outcomes, and the median moves far more than the mean suggests.
The fix is not to pick worse pilot participants. It is to run a second cohort of conscripts before you scale. Assign, do not volunteer. If the capability survives a group that did not ask for it, you have something real.
Substitution two: you swapped curated data for live data
Almost every pilot quietly cleans its inputs. Someone dedupes the account list, fixes the field mapping, pulls the transcripts that actually recorded. That work is invisible in the results and absent at scale.
In production the agent hits the account with three owners, the opportunity with a close date from 2023, the contact record with a job title that has not been true since the reorg. Output quality does not degrade gracefully. It degrades in a way that destroys trust faster than the wins accumulate.
A pilot measures the capability. Production measures your data.
Before scaling, run the capability against an unfiltered slice of live records and count the failures by cause. If more than a small fraction trace to data quality rather than model quality, you are not buying an AI project. You are buying a data project with an AI budget attached.
Substitution three: you swapped attention for defaults
During a pilot someone is watching. There is a Slack channel, a weekly check-in, a person whose visible job is making the thing work. That attention is a real input, and it is never funded at scale.
In production the workflow competes with everything else in the seller's day. Whatever is default, embedded, and unavoidable gets used. Whatever requires a decision to open gets abandoned in about three weeks.
This is why placement beats capability. A mediocre summary that appears inside the record the rep already opens will outperform an excellent one that lives behind a separate login. Design for the moment of work, not the moment of demo.
Substitution four: you swapped a narrow use case for a broad one
Pilots succeed because they are scoped. One segment, one motion, one clearly bounded job. Rollouts fail because the business case required breadth to justify the spend, so the same capability gets pointed at six motions with different data shapes and different definitions of good.
The honest move is to scale the use case, not the license. Take the narrow job that worked, extend it to the next adjacent segment, and measure again before widening. Slower on the slide. Faster in reality.
What a production-grade pilot looks like
- Two cohorts. Volunteers and conscripts, measured separately. Report the conscript number to leadership.
- Unfiltered data. No pre-cleaning. Log every failure and classify it as model, data, or process before you decide what to fix.
- A named owner with a budget line. If nobody owns the workflow after the pilot ends, the workflow ends when the pilot does.
- A baseline captured before day one. Cycle time, conversion, hours spent, quality score. Without it every result is an anecdote.
- Written kill criteria. The number that, if unmet by a date, ends the program. Agreed before anyone falls in love with the demo.
None of this is exotic. It is the difference between running an experiment and running a showcase. Most teams are running a showcase and calling it an experiment.
The question to ask in the next steering meeting
Not did the pilot work. Ask what did we remove from the pilot that production will put back. Cleaned data. Volunteer motivation. Dedicated attention. Narrow scope. Every one of those is a variable, and every one of them is going away.
Teams that answer that question honestly ship fewer AI programs and keep more of them. That is the trade worth making.
Take it to the room
The short list this issue leaves you with
Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.
Checklist diagram summarising Why Your AI Pilot Worked and Your Rollout Did Not: Two cohorts; Unfiltered data; A named owner with a budget line; A baseline captured before day one; Written kill criteria.Frequently asked questions
- Why do AI pilots succeed but rollouts fail?
- Pilots run on volunteers, curated data, dedicated attention, and a narrow use case. Production removes all four. The capability is unchanged, but the conditions that produced the result are gone, so the measured lift disappears.
- How do you design an AI pilot that predicts production results?
- Run two cohorts, volunteers and assigned participants, and report the assigned cohort's numbers. Use unfiltered live data, capture a baseline before day one, name an owner with an ongoing budget line, and write kill criteria before launch.
- What is the biggest hidden cost of scaling AI in a revenue team?
- Data remediation. Most rollout failures trace to record quality, ownership conflicts, and stale fields rather than model quality. Classify pilot failures by cause before scaling so you know whether you are funding an AI project or a data project.
- How long should an AI pilot run before scaling?
- Long enough to survive a full sales cycle stage and at least one period of competing priorities, typically six to twelve weeks. Short pilots measure novelty. Adoption curves usually break around week three, after the initial attention fades.
Subscribe
Get the next post-mortem, including what got turned off.
Keep reading
Reality Check
Agentforce Pricing Explained: Credits, Licenses, and Total Cost
Agentforce is not one price. It is a stack of editions, entitlements, meters, and platform costs that only resolve into a number once you know which product you are buying. Here is how to work out which one you are looking at, and which question to ask next.
Reality Check
What Are Salesforce Core, Advanced, and Max Editions, and How Many Flex Credits Does Each Include?
On Sep 3, 2026 Salesforce published Core ($195), Advanced ($395), and Max ($550) per user/month for Agentforce Sales and Service, with org-level Flex Credit pools of 500,000 / 1 million / 2.75 million. Credits do not scale per seat. Legacy edition pricing stays for existing customers; Agentforce 1 can move to Max at no extra seat price.
