AI incident response runbook (L5)
When an agent sends the wrong thing to a customer or leaks internal content, the clock is minutes. This is the runbook: detect, contain, notify, correct, review, with named owners and pre written customer language.
The steps
- 01
Define what counts as an AI incident
Tool: Manual
Wrong customer output, data exposure, unauthorized action, or sustained quality failure. Ambiguity here means incidents get argued about instead of contained. Owner: security. DoD: incident definitions published with severity tiers.
- 02
Make containment one action
Tool: Manual
A documented kill switch per agent that any on call responder can execute without engineering. If containment requires a code change, you do not have containment. Pitfall: pause controls that only an owner can reach during business hours. DoD: kill switch tested by a non owner.
- 03
Pre write customer notification
Tool: Manual
Approved language for the common cases, cleared by legal in advance. Writing notification copy during an incident costs hours you do not have. Owner: legal plus comms. DoD: templates approved for the top three scenarios.
- 04
Review with a written postmortem
Tool: Notion
Blameless postmortem within five days covering detection time, containment time, and the guardrail that was missing. Publish internally. Owner: incident owner. DoD: postmortem published with at least one guardrail change.
Tools in this playbook
- Manual
- Notion
Next playbooks
Unfamiliar terms are defined in the AI and Revenue Dictionary. Related frameworks live in the framework library.
