← All packs

Blameless postmortems: a checklist for the three ways they quietly fail

"Blameless" gets treated as a tone problem — strip the names out, use passive voice, done. That's not what makes a postmortem useful, and it's not usually what makes one fail. A postmortem can name no one and still be worthless, because it stopped at the first plausible cause, described the incident as if everyone knew at 2am what's obvious now, or ended in an action list that wouldn't have prevented the thing it's attached to. Three checks catch most of that before the document ships.

1. Trigger mistaken for root cause

The trigger is the thing that started the clock: a deploy, a traffic spike, a dependency outage. The root cause is why the system was exposed to that trigger at all. "The 14:02 deploy caused the outage" is a trigger. "The deploy could ship a config value with no validation, and staging uses a different config path so nothing would have caught it" is a root cause. Most first drafts stop at the trigger because it's the easy, well-evidenced part — it's right there in the deploy log — while the root cause requires asking "why did the system let that happen" at least once more.

Check: for every stated cause, ask "why was the system vulnerable to this" one more time. If the answer is another fact about this one incident, keep asking. Stop only when the answer is a standing property of the system (a missing validation, a monitoring gap, an untested assumption) that existed before this incident and will cause the next one too if left alone. If there's more than one contributing factor, list all of them — don't collapse to a single tidy cause because it makes the document shorter.

2. Hindsight bias in the timeline

Once you know how it ended, every early signal looks like it should have been obvious. A postmortem written after the fact will naturally describe the 14:05 alert as "the warning sign that was missed," when in real time it was one of forty alerts that day and looked identical to thirty-nine false positives. This isn't a blame problem, it's an accuracy problem — a timeline that silently imports hindsight teaches the wrong lesson (the real gap: this alert type is indistinguishable from noise) and replaces it with a fake one (the fake lesson: someone should have noticed sooner).

Check: for each timeline entry, write only what was observable and known at that moment, not what it turned out to mean. If an alert fired and was dismissed, state what else was true at that time that made dismissing it reasonable (alert volume, prior false-positive rate, competing incidents) rather than just noting it was dismissed. If the honest answer is "there was no way to distinguish this from normal noise at the time," that's a finding about the alert's signal quality, not about the person who saw it.

3. Orphaned follow-up actions

"Improve monitoring" and "add more tests" show up in almost every postmortem and almost never get done, because they aren't tied to anything specific enough to act on or to check off. The other failure in the same family is the reverse: an action item that's concrete but doesn't map to any root cause in the document — busywork that made it into the list because it sounded relevant, not because the analysis called for it.

Check: for every root cause identified in step 1, confirm at least one follow-up action addresses it specifically enough that a reviewer could verify it happened (a named alert added, a specific validation added to a specific field, a staging config change) — not a vague direction. Then check the reverse: for every follow-up action listed, trace it back to a root cause in the document. An action with no corresponding cause is a guess about what might help, not a lesson from this incident, and should either be justified or cut.

Why this order matters

Trigger-vs-root-cause comes first because everything downstream depends on it — a timeline and action list built around the wrong cause are precise about the wrong thing. Hindsight bias comes second because it corrupts the evidence the root-cause analysis relies on; a timeline that already assumes the answer will bias you toward whatever cause it's hinting at. Orphaned actions come last because they're a bookkeeping check on the first two — every action should trace to a cause, and every cause should have an action, which is only checkable once both are written down correctly.

None of this requires an LLM to catch — it's a "why" asked one more time, an honest timeline, and a two-way trace between causes and actions. Where a model helps is doing this fast across a messy pile of raw material — logs, alert timestamps, chat messages, deploy history — reconstructing a timestamped timeline, flagging where it's drifting into hindsight language, and checking that every action item traces back to a stated cause — which is what the incident-postmortem skill in the Claude Code Power Pack does. If this incident also involved a migration, see the companion guide on safe database migrations; if a dependency bump was in the mix, see the one on dependency upgrade risk; if the postmortem's "root cause" needs re-checking, see the one on root-cause tracing.