Revenue Finance
Scoring an AI Initiative With Proof
Score one AI initiative out of five on the PROOF checks and return a scale, fix, or kill verdict that survives a CFO review.
Where it came from
- Source: Report framework library
Why it was chosen
Encodes a Revenue AI Report framework so the agent applies the published method instead of improvising one.
How to use it
- 1.Copy the SKILL.md text below, or download the raw file.
- 2.Create a folder named exactly scoring-an-ai-initiative-with-proof in your agent's skills directory.
- 3.Save the file inside that folder as SKILL.md.
- 4.Ask the agent one of the trigger requests below.
- 5.Check the output against what you already know before it leaves your desk.
Ask it this
- Run PROOF on our AI SDR platform before the renewal next month.
- The CFO is asking whether this AI spend produced anything. Score it.
- Is this pilot ready to scale to the whole team?
- We have no kill line on this tool. Build the scorecard and tell me the verdict.
Do not use it for
- Our whole AI program reports adoption but nothing moved. Diagnose the pattern.
- Label these initiatives Optimize, Amplify, or Reinvent.
- Write the agenda for next week's revenue meeting.
The SKILL.md file
--- name: scoring-an-ai-initiative-with-proof description: Score one AI initiative out of five on the PROOF checks and return a scale, fix, or kill verdict that survives a CFO review. --- # Scoring An AI Initiative With PROOF Score one AI initiative out of five on the PROOF checks and return a scale, fix, or kill verdict that survives a CFO review. ## When to use this skill - A renewal decision is due on an AI tool and the case for renewal is a demo and an adoption chart. - A pilot is being proposed for expansion to more seats. - The CFO or the board has challenged whether AI spend produced anything. - A quarterly AI review needs every initiative to carry a score and a named owner. - A vendor or an internal team is claiming a result that has been produced once. ## Inputs to collect - Data lineage for the initiative: what data goes in, where it came from, which model processes it, and what prompt or configuration is used. Source: RevOps and the systems owner, not the vendor datasheet. - The record of every production run with its date and its quality result. Source: the tool's own logs or the workflow owner. - The business metric the initiative was bought to move, in the buyer's words, with the before number and the after number. Source: the original business case plus a CRM or finance report. - The name and title of the human who owns that outcome. Source: the org chart. - The written kill criteria: metric, threshold, date, and who can trigger them. Source: the contract or the pilot charter. - Full cost: license, implementation, admin hours, and inference or usage charges. Source: finance system and vendor invoices. ## Process Every check is a yes or no. An initiative that cannot answer all five is not ready to scale (https://www.therevenueaireport.com/frameworks/proof). 1. **Prove where the data came from.** Confirm you know the source of the data, the model, and the inputs. Anonymous data plus a rented model plus an unknown prompt is an unauditable system. Score yes or no. 2. **Run it again.** Confirm the result has been reproduced in production at the same quality. If it worked once, that was a demo. Score yes or no. 3. **Outcome tied to money.** Confirm the metric is a business outcome, not an activity count. Emails sent, meetings booked, and hours saved are inputs. Revenue, retention, gross margin, and cycle time are outcomes. Score yes or no. 4. **Owner has a name.** Confirm a named human owns the outcome, not the tool. Tools do not miss quota. Score yes or no. 5. **Failure mode is clear.** Confirm there is a defined line at which the initiative is killed or replaced. Score yes or no. 6. **Total and decide.** Apply the published bands in Decision rules and write the verdict with the failing checks named. ## Decision rules - Score five out of five as real work. Three or four means fix the failing checks before scaling. Below three, the initiative is Optimization Theater and should be killed or rebuilt (https://www.therevenueaireport.com/frameworks/proof). These bands are published and are used as stated. - Do not partially credit a check. Every check is a yes or a no, because a half-answered lineage question is an unauditable pipeline in practice (https://www.therevenueaireport.com/frameworks/proof). - Reject activity counts on the O check. Ten times more emails is an activity, and it is where the adoption gap hides (https://www.therevenueaireport.com/frameworks/proof). The measured size of that gap in a single sample of GTM leaders is 87 percent of sales teams reporting more selling time against 13 percent reporting higher quota attainment, a 74-point distance (Scale Venture Partners, n=278, https://www.therevenueaireport.com/research/proof-gap). - Treat hours saved as capacity, not return, unless they became measurable output. Capacity is a maybe. - Refuse a tool as an owner. A tool cannot miss a quota, so an initiative whose owner field names a platform fails the second O check (https://www.therevenueaireport.com/frameworks/proof). - Require the kill line to be triggerable by one named person. An initiative with no defined kill line cannot be stopped by evidence, only by budget exhaustion, and the kill line is what turns a bet into a decision (https://www.therevenueaireport.com/frameworks/proof). - Score the data check against real conditions rather than the vendor's claim. Data preparation absorbs most AI project time, and Jonathan Kvarfordt puts it as "A lot of AI projects fail because they don't have good data, good clean data or structured data. 75 percent of the time is spent on preparing data" (https://www.therevenueaireport.com/frameworks/proof). - Pair a low score with the spend-to-outcome ratio when finance is in the room. The Proof Gap is total AI spend in the quarter divided by AI-influenced pipeline in the same quarter, where under 1.0x is defensible, 1.0x to 4.0x is a watch, and over 4.0x on a rolling four-quarter basis triggers a Reversal Ledger review (https://www.therevenueaireport.com/frameworks/proof-gap). - Do not run PROOF on an initiative with under one full quarter of production runtime, matching the Proof Gap exclusion for investments under a quarter of runtime (https://www.therevenueaireport.com/frameworks/proof-gap). Score it as not yet scoreable and set the review date. - Include verification cost in the money check. Refinement work of checking, repairing, and reverifying is about 60 percent of an agentic task's cost and is routinely omitted from business cases (McKinsey, via https://www.therevenueaireport.com/research/spend-vs-attribution). - The framework publishes no minimum number of production runs for the R check. Set that count with the workflow owner before scoring, write it into the scorecard, and use the same count for every initiative in the review so scores stay comparable. ## Output requirements Deliver this exact scorecard, because the review cycle reads the five rows in order. | Check | Pass or fail | Evidence used | Source | |---|---|---|---| | P - Prove where the data came from | | | | | R - Run it again | | | | | O - Outcome tied to money | | | | | O - Owner has a name | | | | | F - Failure mode is clear | | | | Then deliver: the score out of five, the verdict of scale, fix, or kill using the published bands, the named failing checks with the specific fix for each, the named human owner, the kill line in one sentence, and one board-ready sentence with its caveat. ## Verification loop Validate the scorecard before it is presented. 1. Re-read every failed check and confirm the failure is evidenced by a document, a log, or a record, not by an absence of information you did not go looking for. 2. Re-read every passed check and confirm the evidence is specific. A pass supported by "the vendor says so" is a fail on the P check by definition. 3. Confirm the verdict matches the published bands. Five is scale, three or four is fix, below three is kill or rebuild. A verdict that contradicts the band is an override and must be labelled as one, with the reason. 4. Confirm the owner named is a person with a title, and that the kill line names a metric, a threshold, and a date. 5. Fix any failure and repeat checks 1 through 4 across all five rows, because upgrading one check often changes the total and therefore the verdict. Only proceed to present the scorecard when all four checks pass on the same version of the table and every failing row carries a specific fix. If the initiative has under one quarter of production runtime, stop, mark it not yet scoreable, and set a review date instead of producing a score that finance will discount. ## Quality checks - All five checks are answered yes or no, with no blanks and no partial credit. - Every pass carries named evidence from a system of record. - The outcome metric is revenue, retention, gross margin, or cycle time, not an activity count. - The owner is a person, not a tool or a team. - The kill line has a metric, a threshold, a date, and one person who can trigger it. - Full cost includes admin hours and usage charges, not only the invoice. - The verdict follows the published band or is explicitly labelled an override. ## Limitations - PROOF scores whether an initiative is defensible. It does not score whether it is strategically correct, which is what OAR is for. - A five out of five on a small sample over a short window is still a noisy read. Say so when the window is short. - Attribution is genuinely hard. The scorecard separates what can be defended from what cannot, and it does not manufacture certainty. - The framework publishes bands but not run counts or cost thresholds. Those are team-set and a loose setting produces a loose score. ## Example input An AI SDR platform, 14 months in production, 120k dollars per year plus usage charges. Reported result is 10 times more emails sent and 45 percent more meetings booked. Owner field in the program tracker names the platform. No kill criteria in the contract. Data lineage covers CRM contacts but the enrichment source and the prompt configuration are held by the vendor. Illustrative and synthetic, provided to show output shape. ## Example output P, fail. The enrichment source and prompt configuration are not disclosed, so the pipeline cannot be audited when it breaks. R, pass. The platform has run continuously for 14 months and quality has been sampled monthly by the SDR manager. O, outcome, fail. Emails sent and meetings booked are activity counts. No revenue, retention, margin, or cycle-time movement has been measured over the same window. O, owner, fail. The tracker names the platform, not a person. F, fail. No kill criteria exist in the contract or the charter. Score: 1 out of 5. Verdict: kill or rebuild, per the published band that puts anything below three in Optimization Theater (https://www.therevenueaireport.com/frameworks/proof). Fixes if rebuilding: require the vendor to disclose the enrichment source and prompt configuration in writing, name the VP of Sales Development as outcome owner, replace the reported metric with meetings-to-opportunity conversion measured against the prior four quarters, and write a kill line of no conversion improvement by day 90 that the VP can trigger alone. Board sentence: this platform has run reliably for over a year and has never been measured against a revenue outcome, so we are rebuilding the measurement before we defend the renewal, and the decision needs human review at the next finance review. ## Rules of conduct - Write for a Director, VP, or operator. Short sentences. Explain uncommon terms. - Separate facts from assumptions. Never hide uncertainty. - Do not invent numbers, benchmarks, quotes, or customer names. - Do not send messages, change CRM records, or publish anything unless the user explicitly asks. - Flag when a decision needs human review. ## Evidence - https://www.therevenueaireport.com/frameworks/proof - https://www.therevenueaireport.com/frameworks/proof-gap - https://www.therevenueaireport.com/frameworks/optimization-theater - https://www.therevenueaireport.com/frameworks/reversal-ledger - https://www.therevenueaireport.com/research/proof-gap - https://www.therevenueaireport.com/research/spend-vs-attribution - https://www.therevenueaireport.com/research/rollback - https://www.therevenueaireport.com/data/proof-gap-index ## Cite this framework Kvarfordt, Jonathan. "PROOF." The Revenue AI Report. https://www.therevenueaireport.com/frameworks/proof
Common questions
- What does the Scoring an AI Initiative With Proof skill do?
- Score one AI initiative out of five on the PROOF checks and return a scale, fix, or kill verdict that survives a CFO review.
- Where does the Scoring an AI Initiative With Proof skill come from?
- Report framework library. It was written by The Revenue AI Report against a 12 criterion quality rubric and graded in an independent scoring pass.
- Why was the Scoring an AI Initiative With Proof skill chosen for this library?
- Encodes a Revenue AI Report framework so the agent applies the published method instead of improvising one.
- When should the Scoring an AI Initiative With Proof skill not be used?
- Do not use it for: Our whole AI program reports adoption but nothing moved. Diagnose the pattern. Or: Label these initiatives Optimize, Amplify, or Reinvent. Or: Write the agenda for next week's revenue meeting.
- How do I install the Scoring an AI Initiative With Proof SKILL.md file?
- Download the file, create a folder named exactly scoring-an-ai-initiative-with-proof inside your agent's skills directory, and save the file inside it as SKILL.md. The agent loads it when a request matches the description.
Raw file: https://www.therevenueaireport.com/agent-skills/scoring-an-ai-initiative-with-proof/SKILL.md. Plain-language skills with worked examples live in the Skills and Prompts library.
