Uptime monitoring with PagerDuty
A monitor going down triggers a PagerDuty incident, and its recovery resolves it. One incident per outage, with a public status page for everyone who is not on call.
- Triggered02:51Northwind: API is down. timeout at https://api.northwind.example/health
- Resolved03:04Northwind: API recovered after 13 minutes
Setting it up
- In PagerDuty, add an Events API v2 integration to the service that should page, and copy its integration key.
- In StatOSS, open the page's settings, Alerts, choose PagerDuty under "Add a destination", paste the key and save.
- Press "Send a test alert". The test triggers and resolves an event, so the on-call rotation gets a look at it before the first real one.
PagerDuty is on Hobby ($4 a month) and Pro. Email alerts are on every plan, Free included. The pricing page has the rest.
What on-call sees
A monitor going down triggers one incident, keyed to the monitor, with the page, the URL, and the reason in the summary. The recovery resolves that incident, so the two-in-a-row rule and the second location's confirmation mean on-call is not paged for a blip and not paged twice for one outage. Slow alerts trigger and resolve the same way when a monitor has a slow threshold. The public page is the place to send everyone else, and its incident opens by itself when the monitor goes down, so the status page and the pager agree.
When an alert goes out
One alert per change of state: when a monitor goes down, when it turns slow, and when it is back, with the time, the reason, and how long it was out. A monitor counts as down after two failed checks in a row, each confirmed by a second server on another network; a single failed check sends nothing. A repeat interval sends a "still down" notice every so many minutes while an outage lasts, and nothing is sent during a maintenance window. Every kind of monitor alerts the same way: a website, an API, a port, a certificate, a cron job.