Skip to content
Back to skills

Incident Runbook

ASecurity

Production incident response: diagnose → mitigate → resolve, then a blameless postmortem and a repeatable runbook. Stop the impact first, root cause second. Use when production is degraded or broken, and again for the postmortem afterwards.

  • 24 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 3, 2026
ai-agentsgo

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 3, 2026

npx -y skills add crewforth/crewforth --skill incident-runbook --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Incident Runbook?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Incident Runbook
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/crewforth-incident-runbook/badge)](https://www.skillsdirectory.com/skills/crewforth-incident-runbook)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: incident-runbook
description: |
  Production incident response: diagnose → mitigate → resolve, then a blameless postmortem and a repeatable
  runbook. Stop the impact first, root cause second.
  Use when production is degraded or broken, and again for the postmortem afterwards.
---

# Incident Response & Runbook

<!-- routing-eval reads the next line; why it sits in the body: AGENT_TEMPLATE.md -->
Trigger phrases: "incident", "incident response", "runbook", "postmortem", "root cause", "outage", "production incident", "post-incident"

Two modes: **live incident** (what to do right now) and **aftermath** (postmortem + runbook). Priority: stopping
user impact > finding the root cause. No panic, one ordered step at a time.

## Live incident — sequence
1. **Acknowledge & classify** — what is the impact (who, how much), severity (SEV1 full outage … SEV3 minor).
2. **Mitigate the impact FIRST** — rollback, turn off a feature flag, shift traffic, scale up. Without waiting on the root cause.
3. **Single coordinator** — it is clear who decides; communication goes through one channel.
4. **Diagnose** — last change? (deploy/migration/config) narrow it down with logs+metrics+traces (observability).
5. **Resolve** — the smallest safe fix; then verify (health check).
6. **Close** — confirm the impact is over; note the timeline (a postmortem input).

## Mitigation reflexes
- Last deploy suspect → **rollback** (deploy revert).
- Suspect feature → turn off the **feature flag**.
- After a destructive migration → restore from backup (db-migration).
- Dependency/service down → circuit breaker / graceful degradation.

## After the incident

Blameless postmortem + producing a durable runbook: **`references/postmortem.md`**.

## Invariant rules
1. **Stop the impact, then understand** — the root cause does not hold up the resolution.
2. **Blameless culture** — the postmortem questions the system, not the person.
3. **Actions are owned + dated** — no "we'll look at it later".
4. **The runbook is executable** — real commands/steps, not wishes.
5. **Make learning permanent** — the lesson goes into an adr/runbook/monitoring, it does not get lost.

Files in this skill

  • SKILL.md2.2 KB
  • references/postmortem.md885 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…