Evaluate AIIntermediateAbout 45 minutesv1.0.0Last reviewed 2026-09-04

Review an AI Pilot

A go, fix, or stop call on an AI pilot based on the criteria you set before it started.

Who this helps

Executive and Founder, Revenue Operations, Sales Leader, GTM Engineering.

When to use it

  • At the end of any pilot period.
  • When a pilot is being extended a second time without a decision.

Information you need first

  • The success criteria set before the pilot started
  • Pilot results: usage, output, quality, cost
  • What the vendor or internal team promised

Quick Prompt

Best for one task. Copy it, add your information, and run it in your AI assistant.

You are reviewing an AI pilot with no stake in the outcome. Criteria set before the pilot: [paste]. Results: [paste usage, output, quality, cost]. What was promised: [paste]. Give me a go, fix, or stop call. Judge results against the original criteria, not against criteria invented after the fact. Where results are ambiguous, say what would resolve them and how long that takes. If the pilot never had real criteria, say that first and stop the review until they exist.

Full SKILL.md preview

Best for repeatable work. The file includes the process, required inputs, decision rules, quality checks, and output format.

---
name: review-ai-pilot
description: A go, fix, or stop call on an AI pilot based on the criteria you set before it started.
license: MIT
metadata:
  author: The Revenue AI Report
  version: 1.0.0
  last-reviewed: 2026-09-04
  source: https://www.therevenueaireport.com/skills/review-ai-pilot
---

# Review an AI Pilot

A go, fix, or stop call on an AI pilot based on the criteria you set before it started.

## When to use this skill

- At the end of any pilot period.
- When a pilot is being extended a second time without a decision.

## Inputs to collect

- The success criteria set before the pilot started
- Pilot results: usage, output, quality, cost
- What the vendor or internal team promised

## Process

1. Find the criteria written before the pilot. No criteria means the pilot already failed its purpose.
2. Score results against those criteria only.
3. Make the call: go, fix, or stop. Fix gets one defined change and one deadline.
4. Write down the call and the reasoning where the next reviewer can find it.
5. If stop: record what was learned and what it cost. That record is the value.

## Decision rules

- Criteria written after results arrive do not count.
- A second extension without a decision is a stop, announced late.
- Fix gets one change and one deadline. Two fixes means stop.

## Output requirements

- Go, fix, or stop with the reasoning tied to original criteria.
- Ambiguities and what would resolve them.
- A written record either way.

## Quality checks

- Judgment uses pre-pilot criteria.
- The call is one word: go, fix, or stop.
- The reasoning is written down.

## Limitations

- Pilot groups are often your most enthusiastic users. Results may not survive contact with the full team.
- The AI cannot weigh strategic value that was never stated as a criterion.

## Example input

Criteria: 30 percent faster first drafts, complaint rate under 2 percent, adopted by 80 percent of the team by day 45. Results at day 60: drafts 35 percent faster, complaints 1.1 percent, adoption 41 percent.

## Example output

Call: fix. Two of three criteria passed, but 41 percent adoption at day 60 means the tool fits some workflows and not others. One change: identify the two workflows where the 41 percent cluster and restrict the rollout to them. Deadline: 30 days, then a go or stop on the narrowed scope. Do not extend twice.

## Review checklist

- Pre-pilot criteria in hand?
- One-word call made?
- Record written for the next reviewer?

## Works with

- Playbook: Make conversation data do work (L3) (sales, L3) https://www.therevenueaireport.com/playbooks/conversation-data-to-work-l3
- Playbook: Turn dormant seats into one shipped workflow (L2) (enablement, L2) https://www.therevenueaireport.com/playbooks/dormant-seats-workflow-l2
- Playbook: Borrow the engineering harness for revenue (L3) (revops, L3) https://www.therevenueaireport.com/playbooks/engineering-harness-for-revenue-l3
- Tool: Spellbook (AI Agents & Workflow) https://www.therevenueaireport.com/tools/spellbook
- Tool: Writer AI Agents (AI Agents & Workflow) https://www.therevenueaireport.com/tools/writer-ai-agents
- Tool: GitLab Duo (Foundation Models & Infrastructure) https://www.therevenueaireport.com/tools/gitlab-duo
- Tool: Braintrust (Foundation Models & Infrastructure) https://www.therevenueaireport.com/tools/braintrust-evals
- Tool: ComfyUI (AI Agents & Workflow) https://www.therevenueaireport.com/tools/comfyui

## Rules of conduct

- Write for a Director, VP, or operator. Short sentences. Explain uncommon terms.
- Separate facts from assumptions. Never hide uncertainty.
- Do not invent numbers, benchmarks, quotes, or customer names.
- Do not send messages, change CRM records, or publish anything unless the user explicitly asks.
- Flag when a decision needs human review.

## Evidence

This skill is grounded in The Revenue AI Report research:
- https://www.therevenueaireport.com/research/rollback
- https://www.therevenueaireport.com/research/named-reversals
- Related analysis: https://www.therevenueaireport.com/blog/ai-sdr-kill-criteria-before-you-sign
- Related framework: https://www.therevenueaireport.com/frameworks/reversal-ledger



Source and updates: https://www.therevenueaireport.com/skills/review-ai-pilot

The process

  1. 1.Find the criteria written before the pilot. No criteria means the pilot already failed its purpose.
  2. 2.Score results against those criteria only.
  3. 3.Make the call: go, fix, or stop. Fix gets one defined change and one deadline.
  4. 4.Write down the call and the reasoning where the next reviewer can find it.
  5. 5.If stop: record what was learned and what it cost. That record is the value.

Decision rules

  • Criteria written after results arrive do not count.
  • A second extension without a decision is a stop, announced late.
  • Fix gets one change and one deadline. Two fixes means stop.

What the output should include

  • Go, fix, or stop with the reasoning tied to original criteria.
  • Ambiguities and what would resolve them.
  • A written record either way.

Example input

Criteria: 30 percent faster first drafts, complaint rate under 2 percent, adopted by 80 percent of the team by day 45. Results at day 60: drafts 35 percent faster, complaints 1.1 percent, adoption 41 percent.

Example output

Call: fix. Two of three criteria passed, but 41 percent adoption at day 60 means the tool fits some workflows and not others. One change: identify the two workflows where the 41 percent cluster and restrict the rollout to them. Deadline: 30 days, then a go or stop on the narrowed scope. Do not extend twice.

Review checklist before you trust the output

  • Pre-pilot criteria in hand?
  • One-word call made?
  • Record written for the next reviewer?

Common questions

What does the Review an AI Pilot skill do?
A go, fix, or stop call on an AI pilot based on the criteria you set before it started.
Who is the Review an AI Pilot skill for?
Executive and Founder, Revenue Operations, Sales Leader, GTM Engineering. It sits at the intermediate level and takes about 45 minutes.
What do I need before I start?
Collect these first: The success criteria set before the pilot started; Pilot results: usage, output, quality, cost; What the vendor or internal team promised.
What is the difference between the quick prompt and the SKILL.md file?
The quick prompt is for one task. Copy it, add your information, run it. The SKILL.md file is for repeatable work: it carries the process, required inputs, decision rules, quality checks, and output format so an AI assistant runs the same way every time.
What should I check before trusting the output?
Pre-pilot criteria in hand? One-word call made? Record written for the next reviewer?
Is it free to use?
Yes. Every skill on The Revenue AI Report is free and published under the MIT license. Attribution is welcome, not required.

Limitations

  • Pilot groups are often your most enthusiastic users. Results may not survive contact with the full team.
  • The AI cannot weigh strategic value that was never stated as a criterion.

Works with

Run the skill, then roll it out with a playbook. Vendor links are supporting context, not a recommendation.

The research behind this skill

License: MIT. Version 1.0.0. Last reviewed 2026-09-04. Raw file: https://www.therevenueaireport.com/skills/review-ai-pilot/SKILL.md

Related skills

Share this skill

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Get the Report

The research behind these skills, weekly.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.