Uptime monitoring, explained
A monitor is a request made on a schedule, and a rule for when the answers mean trouble. How often, from where, and when it is an outage.
Northwind is up.
The monitor is responding, 4 failed checks in the last 24 hours.
API
Up for 7 h 7 min99.72% passed
1,438 checks, median 204 ms, 4 failed checks, 1 timeout
95% under 299 ms
- Timed out1 min, 1 check
- HTTP 5023 min, 3 checks
What a check is
The simplest check requests a URL and passes when a response arrives within a timeout with an acceptable status. Everything else is a refinement: expecting a specific status for a URL that should redirect, requiring a keyword in the body so a page that answers 200 with an error message still fails, sending headers for an endpoint behind a token, recording the response time so a slow service is visible before it is a dead one. The result of every check is a row: when, pass or fail, how long, and why.
How often
The interval sets the longest an outage can go unnoticed. At every five minutes an outage is found within five; at every minute, within one. Most services do not need more than a check a minute, and anything faster costs the monitored service requests without buying much. Certificates and domain expiry dates move once a year, so an hourly or daily check is enough for those. The uptime calculator turns an uptime target into minutes, which shows why the interval matters: 99.9% a month is 43 minutes of downtime, and a five-minute interval spends a tenth of that just noticing.
From where
A check from one server reports that server's view of the world. A routing problem between it and your service looks like an outage; an overloaded checker reports slow responses that are its own. The fix is a second opinion: before a failure counts, another server on another network repeats the check, and only a failure both see goes on the record. More locations tell you about regional reachability, which matters for a global consumer site and hardly at all for a B2B API used from three countries.
When it is an outage
One failed check is not an outage. Networks drop packets, servers restart, a deploy takes two seconds to swap. The usual rule is two or three failures in a row before the state changes, and one success to change it back. The failed check should still be visible somewhere, since a pattern of single failures is worth knowing about, but it should not page anyone. On StatOSS the rule is two confirmed failures in a row to go down and one success to come back, and every failed check shows on the page's strip whether or not it changed the state.
Slowness is its own state. A service answering in four seconds is up by the rules above and unusable to the people waiting. A threshold in milliseconds, with the same two-in-a-row rule, catches that with its own alert.
What to check
- The website and the API: the URLs your customers hit, plus a health endpoint that checks the dependencies.
- The certificate and the domain: the two renewals that take everything down when forgotten.
- DNS: the record that quietly changed.
- Scheduled jobs: backups, reports, queue workers, checked by their absence.
- Ports and hosts for the things that do not speak HTTP.
What to do with the results
Alerts go to whoever fixes things: email, a chat channel, a pager for the serious cases. The results themselves belong on a status page, where customers can see the same failed checks you saw and an incident opens by itself when the state changes. Keeping every check rather than a daily summary is what makes the page worth reading after the fact.