StatOSS

How to write an incident post-mortem

A post-mortem is the record of what broke, why, and what changes. Here is a template that fits on one screen, and how to fill each part in.

northwind.statoss.comPublic page

Past incidents

API is down

, resolved after 24 min. Affects API.

  1. Resolved. The rollback is out and the checks pass.
  2. Identified. A config change routed API traffic to a drained pool. Rolling back.
  3. Investigating. API stopped responding to checks (HTTP 502). Opened automatically.

Post-mortem

Summary

A config change at 11:18 sent API traffic to a pool that had been drained for maintenance. Requests failed with 502 for 26 minutes.

Detection

The API monitor failed at 11:19 and again at 11:20, and opened this incident.

What we are changing

Pool changes go through review. The deploy checks that a pool has healthy hosts before routing to it. The runbook for draining a pool names the routes that use it.

An incident as it reads on the page, with made-up names: opened by the check, updated by a person, resolved when the checks passed, with the post-mortem under it.

The template

Copy it into the incident, or into a document that the incident links to. Write it within a day or two, while the timeline is still in people's heads and the logs are still there.

# <Title: what broke, in plain words>

Date: <date>    Duration: <hh:mm>    Severity: <major / partial / degraded>
Author: <name>  Status: <draft / final>

## Summary
Two or three sentences. What users saw, for how long, and what caused it.

## Impact
Who was affected and how. Requests failed, jobs delayed, data late.
Numbers where you have them: error rate, requests dropped, customers.

## Timeline (UTC)
hh:mm  First failed check / first alert
hh:mm  Someone starts looking
hh:mm  Cause found
hh:mm  Fix applied
hh:mm  Recovery confirmed

## Cause
The chain of events, not the person. What changed, why it was
allowed to change, and why the change had this effect.

## Detection
How it was noticed. Would it have been noticed sooner, and how?

## Recovery
What was done to bring it back, and whether that is the right
fix or a stopgap.

## What we are changing
- <action, owner, date>
- <action, owner, date>

## What went well
The alert fired, the runbook worked, the rollback was quick.

Filling it in

  • Summary. Written last, read first. A customer should be able to stop after it.
  • Impact. Numbers over adjectives. "Sign-in failed for 38 minutes for everyone" beats "some users had trouble signing in".
  • Timeline. In UTC, from the first failed check rather than the first alert, so the gap between the two is visible. A status page that keeps every check gives you the first line for free.
  • Cause. Ask why until the answer is a property of the system: not "the config was wrong" but "nothing checks the config before it is applied".
  • Detection. The question is whether a customer noticed before you did. If they did, the first change on the list is a monitor.
  • What we are changing. Each line has an owner and a date. Three good ones beat ten.
  • What went well. The things that worked are the things to keep when the next change is made.

Blameless means the system

A person made the change that caused the outage; that is nearly always true and nearly never useful. The question a post-mortem answers is why the system let one change have that effect, and what would have caught it. Write the timeline with roles rather than names, and describe decisions with what was known at the time. People who are not afraid of the document tell the truth in it.

Publishing it

Put it under the incident on the status page, where the customers who were affected will look for it. On StatOSS a resolved incident takes a post-mortem from the Incidents tab and shows it on the page; subscribers get the update. Keep the internal version longer if you must, but the public one should carry the cause and the changes, not just an apology. The incidents docs have the mechanics.