AI incident response runbook (L5)

When an agent sends the wrong thing to a customer or leaks internal content, the clock is minutes. This is the runbook: detect, contain, notify, correct, review, with named owners and pre written customer language.

WORKFLOW1Define what counts as an …I incidentManual2Make containment one acti…nManual3Pre write customer notifi…ationManual4Review with a written pos…mortemNotion
4 steps, in order, with the tool that owns each one.
Adoption ladderSix levels from Starter to Rebuilt. This item sits at level 5.L1 StarterOne tool, no workflow changeL2 AssistedAI drafts, humans approveL3 IntegratedWired into CRM and SlackL4 OrchestratedMulti-step, owned, measuredL5 AutonomousAgent runs, human auditsL6 RebuiltThe process itself changes
This playbook belongs at L5 Autonomous. Running it above your level is how pilots stall.
Measures of successmean time to contain; incidents detected by monitoring versus customers; postmortem actions closedPROVE IT WORKEDmean time to containincidents detected by monitoring versuscustomerspostmortem actions closed

The steps

  1. 01

    Define what counts as an AI incident

    Tool: Manual

    Wrong customer output, data exposure, unauthorized action, or sustained quality failure. Ambiguity here means incidents get argued about instead of contained. Owner: security. DoD: incident definitions published with severity tiers.

  2. 02

    Make containment one action

    Tool: Manual

    A documented kill switch per agent that any on call responder can execute without engineering. If containment requires a code change, you do not have containment. Pitfall: pause controls that only an owner can reach during business hours. DoD: kill switch tested by a non owner.

  3. 03

    Pre write customer notification

    Tool: Manual

    Approved language for the common cases, cleared by legal in advance. Writing notification copy during an incident costs hours you do not have. Owner: legal plus comms. DoD: templates approved for the top three scenarios.

  4. 04

    Review with a written postmortem

    Tool: Notion

    Blameless postmortem within five days covering detection time, containment time, and the guardrail that was missing. Publish internally. Owner: incident owner. DoD: postmortem published with at least one guardrail change.

Tools in this playbook

Next playbooks

Unfamiliar terms are defined in the AI and Revenue Dictionary. Related frameworks live in the framework library.

Share this playbook

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.