---
name: measuring-the-proof-gap
description: Produce one organization-level quarterly ratio of AI spend to verified AI-influenced pipeline that survives finance scrutiny, and read it against the published thresholds.
---

# Measuring The Proof Gap

Produce one organization-level quarterly ratio of AI spend to verified AI-influenced pipeline that survives finance scrutiny, and read it against the published thresholds.

## When to use this skill

- A board or CFO asks whether the AI investment produced revenue, and the answer has to be one number.
- Consolidating AI vendor contracts before a renewal cycle and needing an org-level baseline first.
- Building the rolling four-quarter view for a board pack, because single-quarter reads are noisy per the Proof Gap Index methodology (https://www.therevenueaireport.com/methodology/proof-gap-index).
- Comparing your own quarterly ratio to the panel median for your seat and size band.
- A seat leader claims AI moved pipeline and you need the incrementality test applied.

This skill covers the organization-level quarterly ratio only. Payback on a single tool belongs to measure-ai-roi (https://www.therevenueaireport.com/skills/measure-ai-roi).

## Inputs to collect

- Total AI platform contract spend for the quarter, from the finance system or GL, covering AI platform contracts such as Agentforce, HubSpot AI, and Salesforce Einstein and equivalents (https://www.therevenueaireport.com/methodology/proof-gap-index).
- AI seat licenses across the revenue stack for the quarter, from vendor invoices.
- AI-specific professional services invoiced in the quarter, from vendor invoices.
- Internal headcount hours dedicated to AI enablement, from timesheets or manager estimate, valued at loaded cost per the Index methodology.
- Net-new pipeline created in the same quarter where an AI touch is verifiable in the CRM record, from the CRM, exported at opportunity level with the AI-touch field and create date.
- A comparable no-AI-touch cohort from the same CRM export, needed for the incrementality test.
- Which of the eight revenue seats owns each spend line, using the Index seat list of Sales, Marketing, RevOps, Enablement, Customer Success, Partnerships, Executive, and Revenue Finance (https://www.therevenueaireport.com/data/proof-gap-index).
- The prior three quarters of both numbers, because board reporting uses the rolling four-quarter view.

## Process

1. Assemble Total AI Spend in Quarter. Sum platform contracts, seat licenses, AI-specific professional services, and internal AI enablement hours at loaded cost. Artifact: one spend table with a GL or invoice reference on every row.
2. Assemble AI-Influenced Pipeline in Quarter. Keep only net-new pipeline where an AI touch is verifiable in the CRM record. Artifact: one opportunity-level export with the verifying field named.
3. Apply the incrementality test to every retained opportunity. The published test is whether a comparable cohort without the AI touch would have produced the same pipeline. If yes, it does not count (https://www.therevenueaireport.com/methodology/proof-gap-index). Artifact: a kept-and-dropped list with the reason per drop.
4. Strike the four published exclusions from the pipeline side. Time-savings estimates, self-reported productivity claims, pipeline sourced by humans and merely logged by AI, and renewal or expansion pipeline where AI touched only the paperwork. Artifact: the exclusion log.
5. Calculate the single-quarter ratio. Proof Gap = (Total AI Spend in Quarter) / (AI-Influenced Pipeline in Quarter) (https://www.therevenueaireport.com/frameworks/proof-gap).
6. Calculate the rolling four-quarter ratio from the same four quarters of both numerator and denominator. Artifact: both numbers side by side.
7. Classify against the published thresholds and state the resulting action. Artifact: a one-line verdict.
8. Cut the ratio by seat if the spend table supports it, so the Index comparison is apples to apples. Artifact: a seat table matching the Index schema columns quarter, seat, median_proof_gap, q1, q3, and panel_n.
9. Write the board sentence with its caveat, and name every input you could not verify.

## Decision rules

- Read direction first. The Proof Gap is a spend-to-outcome cost ratio expressed as a multiple, so lower is better. Under 1.0x is defensible, 1.0x to 4.0x is a watch state requiring quarterly review, and over 4.0x on a rolling four-quarter basis triggers a Reversal Ledger review (https://www.therevenueaireport.com/frameworks/proof-gap). An agent reading it as a return multiple will invert the verdict.
- Trigger the Reversal Ledger review only on the rolling four-quarter number, never on a single quarter, because single-quarter results are noisy and the rolling view is what a board should see (https://www.therevenueaireport.com/methodology/proof-gap-index).
- Incrementality is required, not optional. Without it every AI tool claims credit for pipeline the human team would have generated anyway, which is what makes the number defensible to a CFO (https://www.therevenueaireport.com/methodology/proof-gap-index).
- Never substitute time-savings estimates for pipeline verification. Time saved is an input claim, and the published exclusion list rejects it (https://www.therevenueaireport.com/methodology/proof-gap-index).
- Exclude R&D-classified AI spend where the mandate is capability building rather than revenue attribution, and exclude any AI investment with under one full quarter of production runtime (https://www.therevenueaireport.com/frameworks/proof-gap).
- Do not benchmark against a published panel median yet. As of the current dataset the Proof Gap Index has no observations published and ships as a header-only CSV, so only the schema and method are public (https://www.therevenueaireport.com/data/proof-gap-index). Compare your own quarters to each other instead.
- If the AI-touch field does not exist in the CRM, stop and report the gap rather than reconstructing touches from memory. A team that cannot attribute cannot disprove a productivity claim, and a claim that cannot be disproven survives the quarterly review (https://www.therevenueaireport.com/research/proof-gap).
- Count the verification layer as AI spend. McKinsey found refinement work of checking, repairing, and reverifying is about 60 percent of an agentic task's cost, and it is routinely omitted from the business case (https://www.therevenueaireport.com/research/spend-vs-attribution).
- Set no internal threshold beyond the three published bands. If leadership wants a tighter internal trigger, set it by taking your own worst rolling four-quarter reading and moving one band tighter, and label it as a team-set threshold rather than a published one.

## Output requirements

- Both ratios, single-quarter and rolling four-quarter, each to two decimal places.
- The threshold band and the action it triggers.
- The spend table and the pipeline table with the exclusion log attached.
- The board sentence with its caveat.

Use this table shape.

| Field | Value | Source |
|---|---|---|
| Quarter | Q1 2026 | finance close |
| Total AI spend | $1.2M | GL plus vendor invoices |
| AI-influenced pipeline, incremental | $2.8M | CRM opportunity export |
| Single-quarter Proof Gap | 0.43x | calculated |
| Rolling four-quarter Proof Gap | value | calculated |
| Band | Defensible, under 1.0x | frameworks/proof-gap |
| Action | Report to board, no Ledger review | frameworks/proof-gap |

## Verification loop

1. Validate the numerator. Re-add the spend table against the GL total for AI cost centres. If the two differ by more than the rounding you declared, find the missing line and re-add. Do not proceed while a line is unexplained.
2. Validate the denominator. Re-run the incrementality test on a random sample of retained opportunities. The Report publishes no sample size for this check, so set one and state it: a common working default is ten percent of retained opportunities, floored at twenty opportunities so a small denominator does not produce a one-deal sample. If any retained opportunity fails the comparable-cohort test on the second pass, drop it, recalculate the ratio, and re-sample another batch of the same size.
3. Validate the exclusions. Confirm none of the four published exclusion categories survived into the denominator. If one did, strike it and return to step 2.
4. Validate the rolling window. Confirm all four quarters use the same spend inclusions and the same AI-touch field definition. A definition change mid-window invalidates the rolling number, so restate the earlier quarters or shorten the window and say which you did.
5. Only proceed when the numerator reconciles to finance, the sampled denominator passes incrementality twice, no excluded category remains, and all four quarters use one definition set. Write the board sentence only after that point. If any of the four fails, report the Proof Gap as not calculable this quarter and name the blocking input.

## Quality checks

- Numerator includes internal AI enablement hours at loaded cost, not only invoices.
- Denominator is net-new pipeline only, with a named CRM field proving the AI touch.
- Every excluded item appears in the exclusion log with a reason.
- The verdict cites the published band, not an invented one.
- Both the single-quarter and rolling four-quarter numbers appear together.
- Direction is stated in words, so no reader mistakes a low ratio for a bad result.

## Limitations

- The Proof Gap measures pipeline, not closed revenue, so it is a leading indicator and will move before the P&L does.
- Incrementality by comparable cohort is a judgment call in any org without holdout discipline. This method makes the judgment visible; it does not remove it.
- The Index has no published observations yet, so there is no external median to compare against (https://www.therevenueaireport.com/data/proof-gap-index).
- The ratio says nothing about capability building. R&D-classified AI spend is excluded by design and needs its own read.
- The research base is largely self-reported executive survey data. Scale Venture Partners with Benchmarkit measured adoption and outcome inside the same n=278 sample of GTM leaders, published November 3, 2025, and both sides are self-reported (https://www.therevenueaireport.com/research/proof-gap).

## Example input

A 400-seat sales organization spends $1.2M on AI in Q1 2026 across Agentforce seats, an AI SDR platform, a signal tool, and enablement services. Verified AI-influenced net-new pipeline in the same quarter is $2.8M. Three prior quarters are available on the same definitions. These figures are the framework's own published worked example (https://www.therevenueaireport.com/frameworks/proof-gap).

## Example output

Single-quarter Proof Gap = 1.2 / 2.8 = 0.43x. That is defensible, because under 1.0x is the published defensible band (https://www.therevenueaireport.com/frameworks/proof-gap). Rolling four-quarter Proof Gap is the number the board sees, and it must be calculated on the same four quarters of both inputs before this verdict is reported upward.

Action: report, no Reversal Ledger review, because the trigger is over 4.0x on a rolling four-quarter basis (https://www.therevenueaireport.com/methodology/proof-gap-index).

Board sentence: AI spend of $1.2M in Q1 2026 sits against $2.8M of incremental AI-influenced net-new pipeline for a Proof Gap of 0.43x, inside the defensible band; the caveat is that incrementality rests on a comparable-cohort test rather than a holdout, and $340k of enablement hours are valued at loaded cost rather than invoiced.

Context to hand the CFO alongside it: across every GTM function the reported activity gain outruns the revenue gain. In sales, 87 percent of teams report reps spending more time selling and 13 percent report higher quota attainment inside the same sample, a 74-point distance (https://www.therevenueaireport.com/research/proof-gap). The share of finance leaders who can tie AI spend to a business outcome is 22 percent, from CloudZero with n=260 finance executives including 135 CFOs (https://www.therevenueaireport.com/research/spend-vs-attribution).

## Rules of conduct

- Write for a Director, VP, or operator. Short sentences. Explain uncommon terms.
- Separate facts from assumptions. Never hide uncertainty.
- Do not invent numbers, benchmarks, quotes, or customer names.
- Do not send messages, change CRM records, or publish anything unless the user explicitly asks.
- Flag when a decision needs human review.

## Evidence

- https://www.therevenueaireport.com/frameworks/proof-gap
- https://www.therevenueaireport.com/methodology/proof-gap-index
- https://www.therevenueaireport.com/data/proof-gap-index
- https://www.therevenueaireport.com/research/proof-gap
- https://www.therevenueaireport.com/research/spend-vs-attribution
- https://www.therevenueaireport.com/frameworks/reversal-ledger
- https://www.therevenueaireport.com/skills/measure-ai-roi

## Cite this framework

Kvarfordt, Jonathan. "The Proof Gap." The Revenue AI Report. https://www.therevenueaireport.com/frameworks/proof-gap