What is a status page?
A public page that says whether your service is working right now, what is being done when it is not, and what happened before. What one should hold.
Northwind is up.
All 3 monitors are responding, 2 failed checks in the last 24 hours.
Website
Up for 23 h 58 min100% passed
1,438 checks, median 137 ms, all passed
95% under 199 ms
API
Up for 7 h 6 min99.86% passed
1,438 checks, median 204 ms, 2 timeouts
95% under 299 ms
- Timed out2 min, 2 checks
Docs
Up for 23 h 58 min100% passed
1,438 checks, median 94 ms, all passed
95% under 136 ms
The page in one paragraph
A status page answers the question a customer asks when something looks wrong: is it me, or is it them? It lists the parts of the service, says whether each is up, and shows the recent past, so the reader can tell a blip from an outage. When something is down, it says what is known and when the next update is due. When it is over, it keeps the record. Companies put the page at a separate address (status.example.com) so it stays up when the product does not.
Who reads it
Customers, first: the page is the alternative to a support ticket that says "is it down?". Your own support team, who need to answer that ticket with a link rather than a guess. Engineers on other teams, whose service depends on yours. Buyers doing due diligence, who look at the incident history to see how often things break and how the company talks about it. And search engines: "example down" is a query, and the page you control is the answer you want ranking.
What should be on it
- The current state, as a headline anyone can read at a glance: everything is up, part of it is down, all of it is down.
- The parts, one line each: the website, the API, the dashboard, the mobile app, the thing that sends email. Group them when there are many.
- The recent past for each part. The usual form is a row of daily squares and a percentage. The more useful form is every check with its response time, so a slow hour and a five-minute outage are both visible; that is the form StatOSS draws.
- Open incidents with dated updates, what is affected, and when to expect the next word.
- Planned maintenance, announced ahead, so a deliberate outage is not mistaken for an accident.
- Past incidents with their post-mortems. This is what buyers read.
- A way to subscribe: email, RSS, a webhook, so nobody has to keep refreshing.
Let the checks update the page
A status page that says everything is fine during an outage does more harm than no page at all, because the reader learns it cannot be trusted. So let the checks update the page rather than a person: a monitor that fails opens the incident by itself, and the failed checks are drawn where visitors can see them. A person then adds what the checks cannot know: what broke, what is being done, when it will be fixed. The best practices guide has the rest of the habits.
Getting one
Hosted status pages start at free: StatOSS Free holds one page and one monitor checked every 5 minutes, and Hobby holds 10 monitors checked every minute for $4 a month. Products like Atlassian Statuspage or Instatus sell the page alone and take status from your monitoring tool; the comparisons go through the differences. Open source options let you run the page on your own server; the open source status page guide lists them.