Playbook

AI Supplier Shortlisting: The Nine-Question Diligence Pass That Cuts 40 Vendors to 3

A shortlist is an elimination process, not a research project. Nine questions, five gates, and a scoring sheet that gets a revenue team from a crowded category to three defensible finalists in two weeks.

Jonathan Kvarfordt · Published September 6, 2026 · 12 min read

Why trust this analysis?

The short answer

How do we shortlist AI suppliers without spending a quarter on demos?

Eliminate on disqualifiers before you evaluate on merit. Five gates cut a forty-vendor category to a handful in days: ownership stability, evidence class, integration reality, unit economics, and exit terms. Only then do you run the nine diligence questions against the survivors.

Decision rule

Any vendor that fails a gate is out, regardless of demo quality. Demos test the seller. Gates test the purchase.

Operator action

Run the two-week sequence below with a named owner per gate.

Supporting pages

Last reviewed

Most AI shortlists fail the same way. A team collects forty names, books twelve demos, watches the same slide deck twelve times, and ends up choosing the vendor with the best presenter. The process rewarded sales skill and told you nothing about the purchase.

Shortlisting is elimination. You are not looking for the best vendor. You are removing every vendor you would regret, until the remainder is small enough to test properly.

The argument

How this playbook breaks down

A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.

Contents diagram for AI Supplier Shortlisting: The Nine-Question Diligence Pass That Cuts 40 Vendors to 3, listing the sections: Gate one: ownership stability, Gate two: evidence class, Gate three: integration reality, Gate four: unit economics you can compute, Gate five: exit terms, The nine diligence questions for the survivors.

Gate one: ownership stability

Before capability, ask who owns the company and for how long. Acquisitions and renames reset roadmaps, support models and pricing. A category can turn over hard in a year: engagement platforms folded into execution suites, signal tools absorbed by conferencing platforms, image tools sold twice. We track the ownership state of every tool in the AI tech landscape for exactly this reason.

Eliminate any vendor whose ownership changed in the last twelve months and cannot produce a written roadmap commitment for the product you are buying. Not a verbal one from an account executive.

The gates

Eliminate before you evaluate

Five gates, applied in order, before any demo is booked. Schematic, not a dataset. Source-cited charts live in the research library.

Eliminate before you evaluate. Diagram showing Ownership, Evidence, Integration, Unit economics, Exit.

Gate two: evidence class

Sort every claim into one of four classes and score them differently.

  1. Customer-published. A named customer, on their own channel or on the record, describing a result. Highest weight.
  2. Vendor-published with a named customer. A case study on the vendor's site with a real logo and quote. Useful, discounted.
  3. Third-party aggregate. Review platform scores with visible volume. Directional, not decisive, and thin volume is a finding in itself.
  4. Unattributed. Percentages with no customer, no sample and no date. Zero weight, and a signal about how the vendor treats evidence generally.

This is the proof gap applied at the shortlist stage. If a vendor's entire case rests on class four, they are out before the demo. The measured size of that gap across the category sits in the proof gap research.

Gate three: integration reality

Ask three questions and get the answers from an engineer, not a sales engineer.

  • Which system of record does this write to, and does it write or only read?
  • What permissions, API limits and editions does it require, and what degrades if we withhold one?
  • What is the median time from contract to first production output for a customer of our size, measured in their data rather than in the deployment plan?

Agents inherit your CRM, they do not repair it. If your data cannot support the tool, buying the tool does not fix the data. That sequencing argument is in CRM data readiness for AI agents.

Gate four: unit economics you can compute

Every AI purchase must reduce to a cost per unit of output you already measure: cost per held meeting, per resolved ticket, per qualified opportunity, per renewal saved. If the vendor cannot express their price in your unit, you cannot compare them to anything, including doing nothing.

Include four costs, not one: the licence, the data and enrichment it depends on, the human supervision hours, and the remediation time when output is wrong. The full model is in AI SDR unit economics, and it generalises past SDRs.

On consumption pricing, add one more question: who defines the billable unit and who can audit it. That question is worked in full in who audits the meter.

Gate five: exit terms

Write the exit before you write the business case. Term length, break rights, data export format, and the kill criteria that trigger a review. A vendor who resists kill criteria is telling you what they expect the results to look like.

The nine diligence questions for the survivors

Everything above removes vendors. These nine choose between the ones left.

  1. Show the result on our data. Run the model or the workflow against a held-out sample of our own closed history, with no tuning.
  2. Who at our company owns this after go-live, and what is their weekly job with it?
  3. What does the workflow look like the day after we switch it on, step by step, including the human steps?
  4. What is the human review rate today at your largest comparable customer, and what was it at their go-live?
  5. What is your error profile, and what happens downstream when the system is wrong?
  6. Which of our existing tools does this make redundant, and are we prepared to actually retire them?
  7. What is the twelve-month total cost including implementation, data, and the internal hours we just named?
  8. What are the two most common reasons customers of our size churn from you?
  9. What would have to be true in ninety days for us to expand this, and what would have to be true for us to stop?

Question eight is the tell. A vendor who answers it precisely has watched their own failures. A vendor who says customers do not churn has not been selling long enough to have seen one.

The two-week sequence

  1. Days 1 to 2. Build the long list from the category. Apply gates one and two on public evidence alone. Expect to lose half.
  2. Days 3 to 5. Send gates three, four and five as a written questionnaire. Set a deadline. Non-response is a response.
  3. Days 6 to 8. Score the returns. Cut to three finalists. Publish the scoring sheet internally so the decision survives its author.
  4. Days 9 to 12. Run the nine questions live with each finalist, engineer and owner present, and require the held-out test on your data.
  5. Days 13 to 14. Write the recommendation with kill criteria and the rollback plan already drafted.

Two weeks, three finalists, a written rationale and an exit already specified. That is a shortlist. The rest is theatre.

Related: AI vendor diligence questions · Category durability test · Pipeline truth test · Skill: threading build-buy decisions

Take it to the room

The short list this issue leaves you with

Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.

Checklist diagram summarising AI Supplier Shortlisting: The Nine-Question Diligence Pass That Cuts 40 Vendors to 3: Days 1 to 2; Days 3 to 5; Days 6 to 8; Days 9 to 12; Days 13 to 14.

Frequently asked questions

How many vendors should reach a demo?
Three. Everything above three is a research habit rather than a decision process, and each additional demo costs more internal hours than it removes risk.
What disqualifies a vendor immediately?
Unattributed evidence only, an ownership change within twelve months with no written roadmap commitment, refusal to run a held-out test on your data, or refusal to accept kill criteria in the contract.
Who should own the shortlist?
One named owner per gate, with a single decision owner over all five. Committee scoring without named gate owners produces averages, not decisions.
How do we compare consumption pricing to seat pricing?
Convert both to cost per unit of output you already measure, then ask who defines and audits the billable unit on the consumption side.
What if the category is too new for customer-published evidence?
Then price for that. Shorter term, smaller committed volume, tighter kill criteria. Immaturity is not disqualifying, but it should show up in the contract rather than only in the risk register.
Does this work for build-versus-buy?
Yes. Treat the internal build as a vendor and run it through the same five gates, including the unit economics and the exit terms.

Share this issue

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Subscribe

Get the next operator playbook in your inbox.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.

Keep reading