Optimization Theater: Five Tells That Your AI Dashboard Is Lying to You
The metric improved and the business did not. Five reliable tells, the counter-question for each, and the one-page reporting format that makes theater impossible to sustain.
Jonathan Kvarfordt · Published September 7, 2026 · 11 min read
The short answer
How do we tell whether an AI result is real or theater?
Decision rule
If the metric was invented after the deployment, it is not evidence. It is a narrative device.
Operator action
Apply the five counter-questions below at the next AI review and require the one-page format afterward.
Supporting pages
- The Proof Gap has a measured size the data behind this piece
- Optimization Theater definition
- The Proof Gap definition
Last reviewed
Optimization theater is the practice of improving a measurement rather than an outcome. It is rarely dishonest. It is what happens when a team is asked to show progress on a schedule that does not match the speed of real progress, so the measure bends toward what is available.
The cost is not the wasted spend. It is the credibility damage when a board eventually asks the follow-up question and nobody can answer it. Here are the five tells and the counter-question for each.
The argument
How this reality check breaks down
A map of the sections ahead, in the order the case is made. Schematic, not a dataset. Source-cited charts live in the research library.
Contents diagram for Optimization Theater: Five Tells That Your AI Dashboard Is Lying to You, listing the sections: Tell one: the metric did not exist before the…, Tell two: the numerator moved and the denomin…, Tell three: the cohort changed between the be…, Tell four: activity replaced outcome, Tell five: the comparison window was chosen a…, Why smart teams do this.Tell one: the metric did not exist before the deployment
A new metric appears alongside the new tool. Time saved. Content pieces generated. AI-assisted interactions. None of these existed in last year's operating review, and none of them appear in any plan the board approved.
Counter-question: what was this number twelve months ago? If it cannot be computed, the metric cannot show improvement, only existence.
The tells
Five questions that separate a result from a story
Ask all five at the review. Theater fails at least one. Schematic, not a dataset. Source-cited charts live in the research library.
Five questions that separate a result from a story. Diagram showing What was this number twelve months ago?, What is the denominator?, Were the two groups comparable?, Which revenue line moved?, When was the window agreed, and by whom?.Tell two: the numerator moved and the denominator vanished
Meetings booked went up. Meetings booked per thousand contacts touched went down, and that ratio is not on the slide. Reply rate improved on a list one tenth the size. Conversion improved on a segment quietly narrowed.
Counter-question: what is the denominator, and how did it change over the same period? Rate metrics without a stated base are the single most common form of theater.
Tell three: the cohort changed between the before and the after
The pilot ran with three volunteers who wanted it to work. The comparison group is the rest of the team, who did not. Or the pilot covered enterprise accounts and the baseline covered all segments. Selection does the work the tool was supposed to do.
Counter-question: were the two groups comparable before the deployment, and can you show one pre-period metric where they matched?
Tell four: activity replaced outcome
The report is full of things that happened and empty of things that changed. Emails sent, summaries generated, calls transcribed, seats activated. This is the failure mode a well-known operator post described precisely: definitions you made up, presented instead of revenue you produced.
Counter-question: which line in the revenue plan moved, by how much, and over what period? If the answer requires two steps of inference, treat it as unproven. The discipline is the same one in reporting AI impact to the board.
Tell five: the comparison window was chosen after the fact
Results are shown against the weakest available quarter, or against a period that included a hiring gap, a pricing change or a seasonal trough. Nobody agreed the window in advance because agreeing it in advance would have removed the option to pick.
Counter-question: when was the comparison window agreed, and by whom? Write it into the pilot charter before the pilot starts and this tell disappears permanently.
Why smart teams do this
Three structural pressures, none of them malicious.
- Reporting cadence outruns effect size. Real revenue effects take two to four quarters. Board decks come quarterly. Something has to fill the gap.
- No baseline was captured. Once the before is gone, every honest analyst is stuck constructing a comparison, and constructed comparisons drift toward the flattering one.
- The sponsor is also the evaluator. When the person who approved the spend also grades it, incentives do the rest without anyone deciding to mislead.
The structural fix is to separate the sponsor from the evaluator and to capture the baseline before the tool is switched on. The measured shape of this gap across the category is in the proof gap research.
The one-page format that ends it
One page per AI initiative, same fields every time, reviewed on the same cadence as pipeline.
- The claim. One sentence, naming the business metric.
- The baseline. The value before deployment, with the date it was captured.
- The current value. With the denominator stated.
- The comparison window. Agreed date, agreed by name.
- The counterfactual. What else changed in the period that could explain the movement.
- Cost. Licence, data, human hours, remediation.
- Kill criteria status. Triggered, near, or clear.
Any initiative that cannot fill this page is not ready to be reported. That is not a bureaucratic standard, it is the minimum evidence a sceptic needs, and the sceptic in your board meeting is the one whose opinion decides the next budget.
Related: Reporting AI impact to the board · Pipeline truth test · Optimization theater framework · Skill: detecting optimization theater
Take it to the room
The short list this issue leaves you with
Pulled from the argument above, written so you can read it out in a pipeline or board review. Schematic, not a dataset.
Checklist diagram summarising Optimization Theater: Five Tells That Your AI Dashboard Is Lying to You: The claim; The baseline; The current value; The comparison window; The counterfactual.Frequently asked questions
- What is optimization theater?
- Improving a measurement rather than an outcome. The dashboard moves, the business does not, and the metric usually did not exist before the deployment.
- What is the fastest single check?
- Ask what the number was twelve months ago. If it cannot be computed, the metric cannot demonstrate improvement.
- Is time saved a valid metric?
- Only when the saved hours were reallocated to a named activity and that activity moved a business metric. Otherwise it is capacity that was never banked.
- How do we prevent it structurally?
- Capture the baseline before deployment, agree the comparison window in writing, and separate the person who sponsored the spend from the person who evaluates it.
- What belongs in a board-ready AI page?
- Claim, baseline with capture date, current value with denominator, agreed comparison window, counterfactual, full cost, and kill-criteria status.
- How long before real effects appear?
- Most revenue-level effects take two to four quarters to separate from noise. Reporting cadence should reflect that rather than manufacturing quarterly proof.
Subscribe
Get the next Reality Check before you sign the order form.
Keep reading
Reality Check
Agentforce Pricing Explained: Credits, Licenses, and Total Cost
Agentforce is not one price. It is a stack of editions, entitlements, meters, and platform costs that only resolve into a number once you know which product you are buying. Here is how to work out which one you are looking at, and which question to ask next.
Reality Check
What Are Salesforce Core, Advanced, and Max Editions, and How Many Flex Credits Does Each Include?
On Sep 3, 2026 Salesforce published Core ($195), Advanced ($395), and Max ($550) per user/month for Agentforce Sales and Service, with org-level Flex Credit pools of 500,000 / 1 million / 2.75 million. Credits do not scale per seat. Legacy edition pricing stays for existing customers; Agentforce 1 can move to Max at no extra seat price.
