Playbook

What Kill Criteria Should We Set Before Signing an AI SDR Contract?

Write the fail conditions into the order form before the demo becomes a two-quarter argument. Hybrid 1.9x, reply-rate floors, named redeployment, no headcount-cut ROI story.

Jonathan Kvarfordt · Published August 27, 2026 · 11 min read

Why trust this analysis?

The short answer

Do hybrid AI plus human SDR teams actually outperform?

On efficiency, yes. Bridge Group data shows hybrid teams producing 1.9x meetings per dollar, and 2.4x versus human-only. Raw volume moved from 1,150 to 7,400 touches while reply rates fell from 4.7 percent to 2.9 percent, so volume alone is not the result worth buying.

Evidence

  • The Proof Gap has a measured size Every GTM function adopted AI faster than it produced revenue. One dataset measures both sides in the same sample.
  • What kill criteria should go in an AI SDR contract? Six: meetings per dollar against a documented human baseline, a reply-rate floor with sends capped, named redeployment of recovered hours, a data quality gate with revocable write-back, downstream conversion tracked to opportunity, and an explicit refusal to justify the purchase with headcount reduction.

Supporting pages

Last reviewed

The AI SDR argument almost never gets settled on the merits. It gets settled by attrition: two quarters in, the vendor points at volume, the sales leader points at meetings, finance points at the invoice, and nobody can say what would have counted as failure because nobody wrote it down before the signature.

Kill criteria are the fix, and they belong on the order form, not in a slide. A kill criterion is a named metric, a floor, a measurement window, and a consequence. If you cannot write the consequence, you do not have a criterion. You have a hope.

The argument

How this playbook breaks down

A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.

Contents diagram for What Kill Criteria Should We Set Before Signing an AI SDR Contract?, listing the sections: What the data already says before you pilot, Six kill criteria to write into the contract, The 30, 60, 90, The paragraph to put on the order form.

What the data already says before you pilot

Two things are true at the same time, and holding both is what makes a contract survivable.

Hybrid wins on efficiency, not volume. Bridge Group data covered in AI capacity lift, not headcount cut shows outbound volume moving from 1,150 to 7,400 touches while reply rates fell from 4.7 percent to 2.9 percent. The efficiency result sits underneath that: hybrid teams produced 1.9x meetings per dollar, and 2.4x versus human-only. The volume number is not the win. The per-dollar number is.

Kill criteria

Write the exit before you write the cheque

Each rung is a pre-agreed condition, defined before signature. Schematic, not a dataset. Source-cited charts live in the research library.

Write the exit before you write the cheque. Diagram showing Define the one metric, Set the review window, Name the kill threshold, Agree the exit path, Sign.

Time saved is not revenue. The Gartner findings in the AI reinvestment gap are blunt: 4.8 hours per week returned, and 72 percent of it not reinvested into anything that touches revenue. Teams that named where the hours went saw 2.2x to 3.1x the return of teams that did not, and the ROI split ran 25 percent versus 20 percent depending on whether reinvestment was a written decision or an assumption.

And the headcount-cut version of the ROI story is the one to refuse outright. The capacity-lift issue documents Gartner's finding that teams cutting up to 80 percent of a function saw near-zero ROI gap versus teams that did not cut, Forrester's 55 percent regret rate, Gartner's projection that half of organizations that cut will rehire by 2027, and Robert Half's 29 percent who already reopened the roles. A vendor whose business case is fewer humans is selling you a rehire.

Six kill criteria to write into the contract

1. Meetings per dollar, not meetings

The metric is qualified meetings held per fully loaded dollar of AI SDR spend, measured against your human-only baseline from the two quarters prior. Floor: parity by day 60, and 1.5x by day 90. Hybrid teams in the Bridge Group data hit 1.9x, so a floor below 1.5x by the end of the pilot means your implementation is not reaching the observed range. Consequence for a miss: no renewal, no expansion, and the pilot ends at the term. Write the baseline number into the order form so it cannot be renegotiated later.

2. A reply-rate floor, with volume capped

Volume is the easiest thing for an AI SDR to produce and the least informative. Set a reply-rate floor at no worse than 80 percent of your human baseline, and cap total sends so the vendor cannot buy the meeting number with volume. If reply rate drops below the floor for two consecutive measurement periods, sends pause until messaging is reworked. The industry drift from 4.7 percent to 2.9 percent happened because nobody capped anything.

3. Named redeployment of recovered hours

Before the pilot starts, write down where the recovered hours go: named activity, named owner, named measurement. Multi-threading into open opportunities. Second-call preparation. Customer expansion motions. *Anything except the team will have more time.*** The reinvestment-gap data is the whole justification: 72 percent unreinvested, and a 2.2x to 3.1x return for teams that named the destination. If the redeployment plan does not exist on day zero, the criterion fails at day zero and you should not sign yet.

4. Data quality gate, checked before and during

The AI SDR writes into your CRM. Set the pre-condition: contact and account data must pass your hygiene threshold before sequences start. Then set the running condition: bounce rate, duplicate creation rate, and bad-fit contact rate each carry a ceiling. Breach the ceiling twice and the write-back permission is revoked, not just flagged. An AI SDR that degrades your database is a negative-return purchase even when the meeting number looks fine.

5. Downstream conversion, not top-of-funnel

Meetings held is a mid-funnel metric that can be gamed. Track AI-sourced meetings through to opportunity created and to closed-won, measured against human-sourced meetings from the same period. Floor: AI-sourced meeting-to-opportunity conversion within 25 percent of the human-sourced rate by day 90. If AI-sourced meetings convert at half the rate, the per-dollar number is an illusion and the seller time spent on those meetings is a hidden cost.

6. No headcount-reduction clause, in either direction

State it explicitly in the internal business case: this purchase is not justified by removing SDR headcount, and no headcount decision will be made on pilot data inside the term. The evidence for why is already on the record. Near-zero ROI gap for the teams that cut deeply, 55 percent regret, half projected to rehire by 2027, 29 percent already reopening roles. Refusing the headcount story is what keeps the pilot measurable, because the moment the team believes the tool is here to replace them, every input metric you are measuring becomes political.

The 30, 60, 90

Day 30: instrumentation, not results

Baseline locked and signed by RevOps. All six metrics reporting from a single source. Redeployment plan live with named owners. Data quality gate passing. Nothing about the meeting number is judged at day 30, because a 30-day meeting number is mostly a function of list quality and nothing else. What you are checking is whether you can measure at all. If you cannot, that is the first real finding, and it is usually about your CRM rather than the vendor.

Day 60: parity or explain

Meetings per dollar at parity with the human-only baseline. Reply rate at or above the floor. Bounce and duplicate rates inside the ceiling. Recovered hours visibly landing in the named activity, with the activity's own metric moving. A miss here is not automatically a kill: it is a written explanation from the vendor with a specific change and a date. One documented recovery attempt, not three.

Day 90: the decision

1.5x meetings per dollar or better. Downstream conversion within 25 percent of human-sourced. Data quality clean. Reinvestment measurable. If those hold, expand and keep every criterion in force through the expansion. If they do not, the contract ends at the term with no renegotiation window and the recovered budget goes back to the redeployment activity that was already working.

The paragraph to put on the order form

Ask for this, verbatim, before signature. The pilot term is 90 days. Success is defined as qualified meetings per fully loaded dollar at or above 1.5x the documented human-only baseline, reply rate at or above 80 percent of the human baseline, AI-sourced meeting-to-opportunity conversion within 25 percent of human-sourced, and bounce, duplicate, and bad-fit ceilings unbreached. Metrics are reported weekly from a single mutually agreed source. Failure of any criterion at day 90 terminates at term with no auto-renewal and no early-termination fee. Total sends are capped at the agreed number for the term.

A vendor who has hit 1.9x elsewhere will sign something close to that. A vendor who will not sign any version of it is telling you the ranges in their deck do not reproduce in accounts like yours. That refusal is the cheapest piece of diligence available, and it costs one email.

Set the baseline these criteria are measured against with the AI SDR unit-economics worksheet, then work through the AI SDR hub for the evaluation rubric and the pipeline definitions the contract language depends on.

Take it to the room

The short list this issue leaves you with

Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.

Checklist diagram summarising What Kill Criteria Should We Set Before Signing an AI SDR Contract?: Do hybrid AI plus human SDR teams actually outperform?; What kill criteria should go in an AI SDR contract?; Why isn't outbound volume proof the AI SDR is working?; Should we cut SDR headcount if the AI SDR pilot works?; What is the reinvestment gap and why does it apply here?.

Frequently asked questions

Do hybrid AI plus human SDR teams actually outperform?
On efficiency, yes. Bridge Group data shows hybrid teams producing 1.9x meetings per dollar, and 2.4x versus human-only. Raw volume moved from 1,150 to 7,400 touches while reply rates fell from 4.7 percent to 2.9 percent, so volume alone is not the result worth buying.
What kill criteria should go in an AI SDR contract?
Six: meetings per dollar against a documented human baseline, a reply-rate floor with sends capped, named redeployment of recovered hours, a data quality gate with revocable write-back, downstream conversion tracked to opportunity, and an explicit refusal to justify the purchase with headcount reduction.
Why isn't outbound volume proof the AI SDR is working?
Because volume is the cheapest thing to increase and it drags reply rate down with it. The industry moved from 1,150 to 7,400 touches while replies fell from 4.7 percent to 2.9 percent. Judge meetings per dollar and downstream conversion instead, and cap total sends in the contract.
Should we cut SDR headcount if the AI SDR pilot works?
No, and not on pilot data. Gartner found teams cutting up to 80 percent of a function saw a near-zero ROI gap versus teams that did not cut. Forrester recorded 55 percent regret, Gartner projects half will rehire by 2027, and Robert Half found 29 percent had already reopened roles.
What is the reinvestment gap and why does it apply here?
Gartner found AI returned about 4.8 hours a week while 72 percent of it was never reinvested into revenue-touching work. Teams that named the destination saw 2.2x to 3.1x the return, with an ROI split of 25 percent versus 20 percent. An AI SDR that frees seller time and nothing else lands in that 72 percent.
Why a 90-day break point rather than a full year?
Ninety days is long enough for meetings-per-dollar and downstream conversion to stabilize and short enough that a failing pilot does not become a two-quarter argument. Day 30 checks instrumentation only, day 60 checks parity with one documented recovery attempt, and day 90 is the decision.

Share this issue

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Subscribe

Get the next operator playbook in your inbox.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.

Keep reading