Twenty Questions That Separate an AI Vendor From an AI Demo
Every GTM AI vendor demos well. Diligence is how you find out whether the product works on your data, in your process, at your risk tolerance. These are the questions that change the answer.
Jonathan Kvarfordt · Published June 30, 2026 · 10 min read
The short answer
What questions should you ask an AI vendor before buying?
Evidence
- The same question, four different answers Enterprise AI is paying off for 5% of companies, or 74%, depending on the question asked. None of the four publishers is wrong.
- How do you evaluate AI vendors beyond the demo? Require written answers before the second demo, then run a bounded trial on your messiest segment rather than your cleanest. Define success criteria and a named decision-maker before the trial starts, because trials without a decider convert into subscriptions by default.
Supporting pages
- The same question, four different answers the data behind this piece
- Kill Criteria definition
Last reviewed
The demo is not the product. The demo is the product's best hour, on the vendor's data, driven by the person who built it. You will never see that hour again.
Diligence is how you find out what the other 999 hours look like. Below are the questions that consistently produce a different answer than the sales conversation, grouped by what they protect. Ask them in writing, and treat a vague answer as an answer.
The argument
How this the teardown breaks down
A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.
Contents diagram for Twenty Questions That Separate an AI Vendor From an AI Demo, listing the sections: Questions about the model and the output, Questions about your data, Questions about your process, Questions about the commercial structure, How to run the evaluation so the answers matt…, The meta-signal.Questions about the model and the output
- Which models do you use, and what happens to our workflows when you change them?
- How do you measure output quality today, and can we see the measurement on accounts like ours?
- What is your documented failure rate, and what does a typical failure look like?
- How does the system behave when the input data is missing or contradictory? Does it abstain or guess?
- Can we see three real outputs you consider bad, and what you did about them?
Question five is the one that separates operators from marketers. A vendor with a real quality program can produce bad outputs instantly and talk about them without flinching. A vendor without one will change the subject.
Diligence
The questions to ask before the contract, and again after it
A recurring loop, not a one-time procurement checklist. Schematic, not a dataset. Source-cited charts live in the research library.
The questions to ask before the contract, and again after it. Diagram showing Evidence, Data, Failure, Exit, Roadmap.Questions about your data
- What exactly do you ingest, at what frequency, and what do you retain after we terminate?
- Is our data used to train or improve any model that another customer benefits from? Where is that in the contract, not the FAQ?
- Which subprocessors touch our data, and where do they operate?
- What happens to derived data, embeddings, and summaries on termination?
- Can we run in a mode where nothing leaves our tenant, and what capability do we lose in that mode?
The last one matters more each year. Many vendors offer a private mode that silently disables the features you bought. Ask what breaks, in writing.
A vendor who cannot show you a bad output does not measure quality. They measure demos.
Questions about your process
- Where does this sit in the seller's existing workflow, and what does the rep have to open that they do not open today?
- What is your median time to first value with a customer of our size and our data condition?
- Which of your customers turned this off, and why? We will accept anonymized answers.
- What does implementation actually require from our RevOps team, in hours, by week?
- How does the system handle our exceptions: partner deals, multi-year renewals, whatever is weird about our motion?
The churn question is rarely answered honestly, but the shape of the dodge is informative. A vendor who says nobody has ever turned it off is either very new or not telling you the truth.
Questions about the commercial structure
- What is the unit of pricing, and what happens to our bill if usage doubles?
- What are we contractually entitled to if the quality metric we agree on is not met?
- What is the exit path: data export format, timeline, and cost?
- What is on the roadmap that we are being asked to pay for today but cannot use yet?
- Who owns prompts, configurations, and workflows we build on the platform?
Question sixteen deserves emphasis. If the contract has no quality obligation, you have bought access, not outcomes. That may be fine. It should be a decision, not a discovery.
How to run the evaluation so the answers matter
Send the questions before the second demo, in a document, and require written responses. Score them yourself against what you know about your own environment. Then run a bounded trial on your ugliest segment, not your cleanest, because clean data will make three vendors look identical and ugly data will rank them accurately.
Define what qualified success looks like before the trial starts, and name the person who decides. Trials without a decider become subscriptions by default.
The meta-signal
Across dozens of these evaluations, the strongest predictor of a good outcome is not the model, the funding, or the logo wall. It is how the vendor behaves when you ask a question they cannot answer well. The ones who say we do not measure that yet, and here is what we do measure, tend to ship. The ones who reframe tend to churn.
You are not buying software. You are buying a working relationship with a team that will be wrong sometimes. Evaluate the wrongness.
Take it to the room
The short list this issue leaves you with
Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.
Checklist diagram summarising Twenty Questions That Separate an AI Vendor From an AI Demo: What is the unit of pricing, and what happens to ou…; What are we contractually entitled to if the qualit…; What is the exit path: data export format, timeline…; What is on the roadmap that we are being asked to p…; Who owns prompts, configurations, and workflows we….Frequently asked questions
- What questions should you ask an AI vendor before buying?
- Ask for real examples of bad outputs and how they were handled, documented failure rates, exact data ingestion and retention terms, subprocessor lists, median time to first value for similar customers, customers who turned it off, contractual quality obligations, and the full exit path for your data.
- How do you evaluate AI vendors beyond the demo?
- Require written answers before the second demo, then run a bounded trial on your messiest segment rather than your cleanest. Define success criteria and a named decision-maker before the trial starts, because trials without a decider convert into subscriptions by default.
- What contract terms matter most for GTM AI tools?
- A defined quality obligation with a remedy, clear rules on whether your data trains shared models, ownership of prompts and configurations you build, predictable pricing when usage scales, and a documented export format, timeline, and cost for exit.
- What is the biggest red flag in an AI vendor evaluation?
- An inability to produce examples of bad output. Vendors with a real quality program can show failures immediately and discuss them plainly. Reframing the question usually means quality is measured by demos rather than by measurement.
Subscribe
Get the next teardown in your inbox.
Keep reading
The Teardown
The Signal Stack: What Common Room Buyers Learn After Go-Live
Signal platforms find intent. They do not hand you contacts, and they do not act. Here is the four-layer signal stack, the two costs buyers underestimate, and what the Zoom acquisition changes for renewals.
The Teardown
Ebsta Reviewed: What the Revenue Intelligence Claim Actually Measures
Ebsta sells engagement data as forecast accuracy. Those are two different products. Here is what the review base supports, what the Fullcast acquisition changes, and the eight questions to ask before you swap a forecasting platform for an engagement layer.
