StatOSS

Status page best practices

Ten habits for a page people can rely on during an outage, and five things not to do.

northwind.statoss.comPublic page

Northwind is up.

All 3 monitors are responding, 2 failed checks in the last 24 hours.

Website

Up for 23 h 58 min100% passed

1,438 checks, median 137 ms, all passed

95% under 199 ms

API

Up for 7 h 6 min99.86% passed

1,438 checks, median 204 ms, 2 timeouts

95% under 299 ms

  • Timed out2 min, 2 checks

Docs

Up for 23 h 58 min100% passed

1,438 checks, median 94 ms, all passed

95% under 136 ms

A page that follows the habits, with made-up data: parts named the way customers name them, every check drawn, the failed ones listed with the time and reason.

The habits

  1. Put it somewhere else. A status page on your own infrastructure goes down with it. Use a hosted page on its own domain, and point status.yourdomain at it rather than serving it yourself.
  2. Let the checks open incidents. A person opening an incident twenty minutes into an outage is twenty minutes of a page that said everything was fine. A monitor that fails should open the incident itself; the person adds the words.
  3. Show the failed checks. Daily green squares hide everything shorter than a day. A strip of every check with its response time shows the five-minute outage and the slow afternoon, and the reader can match it to what they saw.
  4. Name the parts the way customers do. "Checkout", "API", "Dashboard", not the internal service names. Group them when there are more than eight.
  5. Update on a schedule during an incident. Say when the next update is due and keep to it, even when the update is "still looking". Silence reads as abandonment.
  6. Announce maintenance ahead. A planned window, posted in advance, with a reminder before it starts and a note when it ends. Checks during the window should not count as downtime.
  7. Publish the post-mortem. Under the incident, within a few days, with the cause and the changes. The template is one screen.
  8. Let people subscribe. Email, RSS, a webhook. Nobody should have to refresh a page to find out it is fixed.
  9. Link to it everywhere the question gets asked. The footer, the docs, the support auto-reply, the error page. A badge in the README for developer tools.
  10. Keep the history. A year of it. The incident list from last spring is what a prospective customer reads, and an empty list looks like a page that was reset.

What not to do

Do not round an outage away: a page that says 100% for a month with a known outage in it is a page nobody will believe next time. Do not wait for confirmation from three teams before opening an incident; open it, then confirm. Do not close an incident the moment the fix ships; wait for the checks to pass and say so. Do not write "a small number of users" when you know the number. And do not make the page pretty at the cost of the information: the strip and the timestamps are the product, the logo is decoration.

StatOSS was built around these habits: automatic incidents, every check on the page, maintenance windows that end themselves, post-mortems under incidents, subscribers, a year of history on Pro. The status page guide starts from the beginning.