Why Does the Same AI Stack Produce Different Results at Two Companies?
AI is power. Everyone has access to the outlet. The gap between two teams running the same model is domain expertise, systems, and process, not the model.
Jonathan Kvarfordt · Published August 18, 2026 · 8 min read
The short answer
If everyone has the same models, where does competitive advantage come from?
Evidence
- What separates the deployments that work The largest gap between AI leaders and everyone else is not technology. It is having decided what to build.
- Why do two teams with the same AI tool get different results? Because the inputs differ. Process clarity, data definitions, who owns the workflow, and how much domain knowledge is embedded in the prompt all vary between teams. Same tool, different inputs, different outputs.
Supporting pages
- What separates the deployments that work the data behind this piece
- The Proof Gap definition
- The Single-Player AI Problem definition
Last reviewed
Two companies buy the same model, the same orchestration layer, and the same seats. One gets a measurable lift. The other gets mediocre output, decides AI is overhyped, and quietly stops renewing. The variable is not the technology. Both had the same technology.
The way I have said it on record for two years: AI is like electricity. Everyone has access to power. What matters is how the power is applied. You can run a city or you can run an electric chair. Same current.
The argument
How this reality check breaks down
A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.
Contents diagram for Why Does the Same AI Stack Produce Different Results at Two Companies?, listing the sections: Why the model stopped being the moat, Where the gap actually opens, The mediocrity loop, What to do about it before the next purchase, The uncomfortable part.The differentiator is not the AI tech. It is the domain expertise of the human applying the AI tech.Jonathan Kvarfordt, on the New Workings podcast
Why the model stopped being the moat
Frontier capability is now purchasable in a browser tab. When capability is democratized, and it should be, the advantage moves to the layer that is not purchasable: how well you understand your own systems, your own process, and your own buyers. That layer is slow to build and impossible to buy in a demo.
Same tool, different result
What separates two teams running the identical AI stack
Access to the model is equal. Everything below it is not. Schematic, not a dataset. Source-cited charts live in the research library.
What separates two teams running the identical AI stack. Diagram showing Throws AI at the problem, Undefined process, Generic prompts, No workflow owner, Usage counted as adoption, Applies domain expertise, Process written down, Expert in the prompt loop, Named owner per workflow, Output quality measured.This is the practical form of the proof gap. Teams that publish a real result almost always did the unglamorous work first. Teams that publish a story usually skipped it and hoped the vendor covered the difference.
Where the gap actually opens
- Definitions. If two managers describe the same pipeline stage differently, the model inherits both descriptions and averages them into nonsense.
- Process. Speed up a bad process and you get more bad output, faster. The AI is doing exactly what you asked.
- Domain knowledge in the prompt. A generic prompt returns a generic answer. Your five-year understanding of why deals in Latin America stall in procurement is the asset. Most teams never write it down, so the model never sees it.
- Ownership. Someone has to be accountable for the workflow, not just for the license. Without that seat, quality drifts and nobody notices until renewal.
- Feedback. If you cannot tell whether output improved, you cannot improve it. Most teams measure usage and call it adoption.
The mediocrity loop
The failure pattern is consistent. A team applies AI casually, gets mediocre output, mediocre output produces mediocre results, and the team concludes AI is not worth the time. The conclusion is wrong and the experience is real. The tool was never the constraint. The input was.
I can name ten companies applying the same underlying model who get very different results. If the technology were the variable, that spread would not exist.
What to do about it before the next purchase
- Write down the process you are about to automate. If it cannot be written in one page, it cannot be handed to an agent.
- Name the outcome first, then work backwards to the tool. Average sale price, ramp time, churn risk. The tool is chosen last, not first.
- Put the domain expert in the prompt loop. The person who knows why deals die should be shaping what the model looks for, not reviewing a summary of it later.
- Set a kill criterion before launch. See kill criteria. An initiative with no failure condition never fails, it just persists.
- Measure output quality, not logins. Usage is an input. Behavior change is the result.
The uncomfortable part
Domain expertise is not a purchase order. It is the accumulated understanding of how your revenue engine actually behaves, and most teams have it distributed across people who were never asked to write it down. The teams pulling ahead did the boring work of documenting systems, process, and definitions, then pointed AI at that. The teams throwing AI at a problem and hoping it resolves are producing the case studies we later file under rollback.
Access is equal. Application is not. That is the whole story.
Take it to the room
The short list this issue leaves you with
Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.
Checklist diagram summarising Why Does the Same AI Stack Produce Different Results at Two Companies?: Write down the process you are about to automate; Name the outcome first, then work backwards to the…; Put the domain expert in the prompt loop; Set a kill criterion before launch; Measure output quality, not logins.Frequently asked questions
- If everyone has the same models, where does competitive advantage come from?
- From domain expertise applied to your own systems and process. The model is a commodity input. The advantage is knowing what to ask it, what good output looks like in your business, and having clean enough definitions and data for the answer to be usable.
- Why do two teams with the same AI tool get different results?
- Because the inputs differ. Process clarity, data definitions, who owns the workflow, and how much domain knowledge is embedded in the prompt all vary between teams. Same tool, different inputs, different outputs.
- Is mediocre AI output a tool problem or a people problem?
- Usually neither in isolation. It is a process problem. Mediocre inputs produce mediocre outputs, which produce mediocre results, which produce the conclusion that AI is not worth the time. The correct read is that the process feeding the tool was never specified.
- What is the first thing to fix before buying more AI?
- Write down the process you intend to automate, in one page, with definitions. If that page cannot be written, an agent cannot execute it reliably either.
Subscribe
Get the next Reality Check before you sign the order form.
Keep reading
Reality Check
Agentforce Pricing Explained: Credits, Licenses, and Total Cost
Agentforce is not one price. It is a stack of editions, entitlements, meters, and platform costs that only resolve into a number once you know which product you are buying. Here is how to work out which one you are looking at, and which question to ask next.
Reality Check
What Are Salesforce Core, Advanced, and Max Editions, and How Many Flex Credits Does Each Include?
On Sep 3, 2026 Salesforce published Core ($195), Advanced ($395), and Max ($550) per user/month for Agentforce Sales and Service, with org-level Flex Credit pools of 500,000 / 1 million / 2.75 million. Credits do not scale per seat. Legacy edition pricing stays for existing customers; Agentforce 1 can move to Max at no extra seat price.
