It is 3:04 a.m. The pager fired. You are half-awake, staring at a dashboard that says p99 latency elevated and a runbook that begins with “Architecture Overview.” You do not need the architecture overview. You need to know whether to roll back, fail over, or go back to sleep. Most runbooks fail at exactly this moment because they were written for a reader who has time to read.

This post is about a runbook structure that works under cognitive load: symptom-first sections, escalation contacts that are actually reachable, and a last-verified date that tells you whether to trust the page at all. The artifact at the end is a template you can adapt this week.

Why runbooks fail at 3 a.m.

Google’s SRE Workbook describes playbooks as containing “high-level instructions on how to respond to automated alerts,” including severity, impact, debugging suggestions, and possible actions. It also notes a hard truth: “Details in playbooks go out of date at the same rate as production environment changes.” For daily releases, that can mean daily drift.

The same chapter documents a real tension inside Google’s own SRE teams: some engineers prefer general playbook entries that change slowly, while others prefer step-by-step instructions to reduce human variability and mean time to repair. The chapter calls this “a contentious topic” and recommends that teams at least agree on the minimal structured details every playbook must contain.

That is the design constraint. A runbook is not a wiki page. It is a decision aid for a tired person under time pressure. Structure it accordingly.

Symptom-first sections, not system-first

Most runbooks are organized by component: Ingestion Service, Auth Service, Billing Worker. That organization mirrors the codebase, which is useful for onboarding and useless at 3 a.m. The on-call engineer does not know which component is at fault. They know what the alert says.

Organize by symptom. Each section title should be something a pager or a human can match against directly:

  • p99 latency > 2s on /v1/orders
  • 5xx rate > 1% on checkout
  • Webhook delivery backlog > 10k
  • Idempotency key collision rate elevated
  • Rate-limit 429s spiking for a single tenant

Inside each section, use the same four-part shape. This is the part that makes the runbook scannable when your hands are cold and your reading comprehension is at 40%.

1. What this means

One or two sentences. Not a paragraph. Example: “Clients are receiving 429s faster than the documented rate-limit contract allows. Either the limit was lowered, a tenant is retrying without backoff, or a shared dependency is throttling us.”

2. First checks (in order)

Numbered commands or dashboard links. Each check should be falsifiable: it either confirms or rules out a hypothesis. Avoid “investigate the logs.” Prefer “run kubectl logs -l app=orders --since=15m | grep 'problem+json' and look for type values.”

3. Likely causes and mitigations

A short table or list. Cause on the left, action on the right. Include the action that is safe to take without approval, and the action that requires escalation. This is where most runbooks get vague, and vagueness is what causes 3 a.m. Slack threads.

4. When to stop and escalate

An explicit condition, not a feeling. “If the 5xx rate has not dropped within 15 minutes of the mitigation, page the payments on-call.” Conditions are checkable. Feelings are not.

Escalation contacts that are actually reachable

Google’s SRE Book lists “clear escalation paths” as one of the most important on-call resources, alongside well-defined incident-management procedures and a blameless postmortem culture. The book also notes that escalation is “generally a principled way to react to serious outages with significant unknown dimensions” — not a failure.

Yet many runbooks list escalation contacts as a name and a Slack handle. At 3 a.m., that is not an escalation path. It is a hope.

A usable escalation block has four fields per contact:

  • Role, not name. Names change; roles are stable. “Payments service owner” beats “Dana.”
  • Primary channel with the actual mechanism: pager rotation name, phone tree, or on-call alias. Not a personal DM.
  • Secondary channel for when the primary does not answer within the response time your team has agreed on. Google’s SRE Book cites typical paging response targets of 5 minutes for user-facing services and 30 minutes for less time-sensitive systems; use whatever your team has committed to, and write it down.
  • Trigger condition: the specific state that justifies waking this person. “If the mitigation in step 3 has been applied and the error rate is still climbing after 15 minutes.”

Two more rules that save time:

  • List the next escalation, not the whole org chart. A runbook that lists seven people is a runbook that lists zero useful people.
  • Include a fallback for “no one answers.” That fallback is usually a secondary on-call rotation or a peer team. Google’s SRE Book describes teams using a secondary rotation as a fall-through for pages the primary misses, and notes that two related teams sometimes serve as each other’s secondary.

The last-verified date is not decoration

Every runbook section should carry a Last verified date and the name of the person who verified it. This is the single highest-leverage field in the document, because it tells the reader how much to trust the rest.

The SRE Workbook is blunt about why this matters: playbook details go stale at the same rate as the production environment. If your team ships daily, a runbook section that has not been touched in six months is a hypothesis, not a procedure.

Two practices make the date meaningful:

  • Verification means execution, not reading. Someone ran the checks and confirmed the commands still work and the dashboards still exist. A read-through is not verification.
  • Staleness has a threshold. Pick one — 90 days is a common choice — and treat sections past it as suspect. The on-call engineer should know to trust the symptom and the escalation block, and to treat the mitigation steps as a starting point rather than gospel.

If you want a forcing function, add a line to your on-call handoff: “Any section you touched this shift, update the verified date.” The SRE Workbook describes handoff emails as a standard part of on-call practice; this is a cheap addition to one you already write.

Where runbooks sit relative to ADRs and RFCs

Runbooks are not the only internal document, and confusing them with decision records is a common source of duplication.

An Architectural Decision Record captures a single architecturally significant decision and its rationale. The adr.github.io project describes ADRs as helping readers “understand the reasons for a chosen architectural decision, along with its trade-offs and consequences,” and notes that a collection of ADRs forms a project’s decision log. ADRs are read when someone asks why is it like this?

RFCs are proposals. They are read before a decision is made, and they are useful precisely because they are not yet authoritative.

Runbooks are procedures. They are read during an incident, and they are useful precisely because they are authoritative right now — which is why the last-verified date matters more here than in either of the other two.

Keep them separate. Link them. A runbook section that says “if you are considering a rollback, see ADR-014 for why the current deploy strategy was chosen” is helpful. A runbook that re-explains the decision is a runbook that will drift out of sync with the ADR.

API contract failures that deserve their own runbook sections

If your team owns an HTTP API, a handful of contract-level failures recur often enough to warrant dedicated symptom sections.

Problem details responses that clients cannot parse

RFC 9457 defines the application/problem+json media type and a small set of members: type, title, detail, instance, and status. The spec says consumers MUST use the type URI as the problem type’s primary identifier, and that clients MUST ignore unrecognized extension members. A runbook section for “clients report unparseable errors” should check whether the response is actually being served as application/problem+json, whether type is an absolute URI, and whether a recent deploy changed an extension member name.

Idempotency key collisions

If your API accepts an idempotency key on mutating requests, a collision or reuse pattern will eventually page someone. The runbook section should state where keys are stored, what the retention window is, and what the safe mitigation is (usually: reject with a clear problem detail, do not silently overwrite).

Rate-limit responses that clients ignore

GitHub’s REST API documentation is a useful reference for what a well-behaved client is expected to do: if a retry-after header is present, do not retry until that many seconds have elapsed; if x-ratelimit-remaining is 0, wait until x-ratelimit-reset; otherwise wait at least a minute. It also warns that continuing to make requests while rate limited may result in the banning of your integration. A runbook section for “429s spiking” should distinguish between a client that is misbehaving and a limit that was legitimately lowered, because the mitigations are different.

Webhook delivery backlog

Webhook failures are usually a queue-depth problem before they are a signature problem. The runbook section should separate the two: check backlog and retry rate first, then check HMAC verification failures, because a signature mismatch on a healthy queue is a different incident than a stalled consumer.

Before and after: the same section, rewritten

Here is a typical runbook section as it is often written, followed by the same content restructured for 3 a.m.

Before

## Orders Service

The orders service handles order creation and retrieval. It depends on the
payments service and the inventory service. It uses PostgreSQL for storage
and Redis for caching. Recent changes include a migration to the new
connection pooler.

If you see errors, check the logs. You may need to restart the service.
Contact the team if it does not resolve.

After

## Symptom: 5xx rate > 1% on POST /v1/orders

Last verified: 2026-09-12 by @priya

### What this means
Clients cannot create orders. Most common causes: connection pool
exhaustion, payments dependency returning 5xx, or a bad deploy.

### First checks (in order)
1. Dashboard: Orders / Error Rate (link)
2. `kubectl logs -l app=orders --since=15m | grep 'problem+json'`
   - Look at the `type` field. `.../upstream-timeout` points at payments.
3. Check connection pool saturation: Orders / DB Pool (link)

### Likely causes and mitigations
| Cause | Safe action | Needs approval |
|---|---|---|
| Pool exhaustion | Scale pooler replicas | — |
| Payments 5xx | Page payments on-call | Rollback orders deploy |
| Bad deploy | — | Rollback to previous release |

### When to stop and escalate
If error rate has not dropped 15 minutes after mitigation, page the
payments service owner (rotation: payments-primary).
If no answer in 5 minutes, page payments-secondary.

The second version is longer. It is also the one you can follow when you cannot remember your own name.

What to leave out

Runbooks accumulate. The SRE Workbook warns that playbooks can “get pulled in many directions” when teams disagree about content, and suggests noticing when playbooks have “accumulated a lot of information beyond these structured details.”

Leave out:

  • Architecture overviews. Link to the diagram instead.
  • Decision rationale. That belongs in an ADR.
  • Onboarding context. That belongs in an onboarding guide.
  • Commands that only work on one engineer’s laptop.
  • Any step that begins with “just” or “simply.”

If a runbook section is a deterministic list of commands that runs every time a particular alert fires, the SRE Workbook recommends implementing automation instead. That is the right long-term move. Until then, the runbook is the automation.

The artifact: a runbook section template

Copy this into your runbook repo. One section per symptom. Keep it boring.

## Symptom: <alert name or human-readable symptom>

Last verified: YYYY-MM-DD by @handle
Owner: <team or role>

### What this means
<One or two sentences. What is broken from the client's perspective.>

### First checks (in order)
1. <Dashboard link or command>
2. <Dashboard link or command>
3. <Dashboard link or command>

### Likely causes and mitigations
| Cause | Safe action | Needs approval |
|---|---|---|
| <cause> | <action> | <action> |
| <cause> | <action> | <action> |

### When to stop and escalate
If <checkable condition>, page <role> via <primary channel>.
If no response within <N> minutes, page <role> via <secondary channel>.

### Related
- ADR: <link>
- Dashboard: <link>
- Postmortem: <link>

Three fields do most of the work: the symptom title, the escalation block, and the last-verified date. Get those right and the rest is refinement.

FAQ

How often should we verify runbook sections?

There is no universal number. The SRE Workbook’s point is that staleness tracks production change rate, so a team shipping daily needs a shorter interval than a team shipping monthly. Pick a threshold your team will actually honor — 90 days is common — and make the date visible so readers can judge for themselves.

Should the runbook live next to the code or in a wiki?

Next to the code, if your review process can enforce updates. A wiki is easier to write and easier to forget. The deciding factor is whether a code change that invalidates a runbook step will trigger a review comment. If it will not, the location does not matter much.

What if we do not have a secondary on-call rotation?

Google’s SRE Book notes that teams without a dedicated secondary sometimes pair with a related team so each serves as the other’s fall-through. If you have no such arrangement, the escalation block should say so explicitly and name the next best contact, even if it is a manager. An honest gap is better than a fictional path.

Do we need a runbook for every alert?

The SRE Workbook describes a practice where a playbook entry is created whenever an alert is created. That is a reasonable default. If an alert fires and no one knows what to do, that is the signal to write the section — not a reason to skip it.

How does this relate to our OpenAPI spec?

The OpenAPI Specification is a formal standard for describing HTTP APIs, and it is the right place for contract details like error schemas and rate-limit headers. The runbook is the right place for what to do when those contracts fail in production. Link them; do not duplicate them.

The next time a runbook section pages you at 3 a.m., check the last-verified date first. If it is stale, trust the symptom and the escalation block, and treat everything else as a suggestion. Then update the date before you go back to bed.