Checks
A monitor is one thing StatOSS checks: a URL, a port, a DNS name, a host, a certificate, a domain, a job that pings back, or a component with no check.
Northwind is up.
All 2 monitors are responding, 2 failed checks in the last 24 hours.
API
Up for 7 h 6 min99.86% passed
1,438 checks, median 204 ms, 2 timeouts
95% under 299 ms
- Timed out2 min, 2 checks
Nightly backup
Up for 23 h 58 min100% passed
1,438 checks, all passed
HTTP
StatOSS requests the URL every minute on the paid plans and every 5 minutes on Free, with a ten-second timeout, from Aarhus and Copenhagen in Denmark, Falkenstein in Germany, Ashburn in the USA, Singapore, and Sydney in Australia. Hobby and Pro can add 7 more, from $4 a month. A check passes when the response arrives in time with a 2xx status. If a monitor has an expected status, the response must match it exactly; that is how a URL that is meant to redirect, or to answer 401, can be checked. Anything else fails: a timeout, a connection or TLS error, or an unexpected status. Loopback, private and link-local addresses are refused, on every redirect hop, since the checks run from our servers.
Under Request options a monitor can use another method (HEAD, POST, PUT, PATCH or DELETE), send headers such as an Authorization token, send a body, and require a keyword: the response must contain the text, or must not. A missing or unwanted keyword fails the check like a bad status would. Bodies are read up to one megabyte, and a keyword past that counts as missing.
The response time runs from looking the host up to the first byte of the response: the lookup, the connection, the TLS handshake and the wait for the server, each shown on its own in the checks behind a bar. The body is read after the clock stops. A host with several addresses is tried the way a browser does: the next one starts after 250 ms without dropping the first, and the first to connect is used.
Other types
The type is picked when a monitor is added and cannot change afterwards, because the history would not mean the same thing. All types share the rules below, the second opinion, the alerts and the incidents.
- TCP port: opens a connection to a host and port and closes it again. The response time is the lookup and the time to connect.
- DNS: asks for one record type (A, AAAA, CNAME, MX, TXT or NS) and passes when at least one answer comes back. Set text the answer must contain to catch a record that changed.
- Ping: one ICMP echo. The response time is the round trip, without the lookup.
- Certificate: a TLS handshake with the host, then a look at the certificate. It fails when the certificate is invalid for the name, expired, or expires within the number of days you set (14 unless changed). Checked once an hour. The monitor shows the date the certificate is valid until, and the API has it as
expiresAt. - Domain expiry: asks the registry through RDAP when the domain expires and fails inside your warning window (30 days unless changed). Asked a few times a day. A registry that publishes no date passes. The monitor shows the date the registration runs until, and the API has it as
expiresAt. - Heartbeat: the direction is reversed. The card shows a URL; your cron job or backup script requests it at the end of every run, with curl, wget or anything else, any method. Set how often a ping is expected and a grace period; once period plus grace pass without one, the next two checks in a row find it missing and the monitor goes down, the same two-in-a-row rule as everything else. The next ping brings it back at once. Before the first ping the clock runs from when the monitor was made.
- Component: no check at all. A part of the product we cannot reach, such as a mobile app or an office network. Its state is set by hand on the monitors tab, and an incident that names it marks it too. Components have no strip and no uptime figure, and they never count against the plan's monitors.
Down, up and slow
A monitor counts as down after two failed checks in a row and as back after one success. Each change sends one alert and opens or resolves an incident on the page. A single failed check still shows on the strip with its reason; it just does not change the state.
A monitor with a response time (HTTP, TCP, DNS, ping) can have a threshold in milliseconds, from 1 to 9999: a check that takes longer than the ten-second timeout is a failure, not a slow success. Two successful responses in a row over it mark the monitor slow, one under it clears that, and both changes are alerted like an outage. The page draws the threshold as a dashed line and colours a bar indigo when it reaches over the line or half its checks were slow. A change to the threshold applies to checks from then on. Slowness never counts against uptime.
On Hobby and Pro a monitor can have a threshold of its own in some regions, under Slow above, by region: a site in Europe can be slow above 300 ms there and above 800 ms from North America. A check from that region is judged against its number, every other check against the one above. When the regions that count toward response time have different thresholds, the strip draws no line, and a bar is indigo when half its checks were slow.
Saved monitors have a Check now button in the dashboard that runs one check straight away and records it like any other.
The second opinion
Before a failed check counts, it is checked again in the same region: from the region's other place, or from the same place a few seconds later when it is the only one there. If that gets through, the check is recorded as passed. Only a failure both see counts toward the two-in-a-row rule. A slow response is checked again the same way, by a place whose speed counts, and is slow only when both were; the reading taken on turn is the one kept. If no answer comes within 20 seconds, the check counts as it was seen. Heartbeats and components need no second opinion; nothing is requested for them.
A place alone in its region that fails on its second look too asks two sites that are always up, Cloudflare's and Google's. When it cannot reach those either, the fault is its own network: the check is not counted, and neither is that region's verdict in the run. An outage is dated from its first failed check, not the second.
The checks behind a bar
Every bar on a strip opens. Click it, or move to it with the arrow keys and press Enter, and a drawer on the right lists the checks behind it: the time of each, the status and response time, and what the second opinion said when it was asked. On pages that show regions it also says where each check ran, what the other regions saw when they were asked, and gives each region's checks and median. Above the list are the totals for the bar, the median, the fastest and slowest response, and any deploy markers inside it. Each check says where its time went (the lookup, the connection, the TLS handshake, the first byte), a failed one how far it got, and when a second place got through after the first failed, what the first saw. In the dashboard it also gives the address each check reached. These are the readings as taken, before any place is levelled to its region's fastest. A bar that holds more than 300 checks, a day on the 90-day range, lists the ones that did not pass cleanly. Bars older than the raw checks (14 days, 7 on Free) show their totals only, and a bar that straddles that age says so. Visitors see a failure in the page's words, such as Timed out; the stored error, which can name a host or a port, is shown in the dashboard only. The drawer sits behind the page's password when there is one.
Regions
Every check location belongs to a region: Europe, North America, South America, Africa, Asia or Oceania. On Hobby and Pro the regions take turns, one check a run, and a region's places take turns inside it. When a check fails, it is checked again in its region and every other region checks at the same moment. While any region has failed lately, every region checks every run.
A region is down after two failed runs there. A monitor is down while a region that counts is down, and says so: down from North America, or down everywhere. Europe cannot overrule North America. Its alerts and its automatic incident name the regions, and a new alert goes out when the outage spreads to other regions or leaves some. A place that fails nearly every check at once while other places do not is set aside for ten minutes. On Free, checks run from Europe.
Hobby and Pro can add places under Account, Check locations: London, Paris, San Jose, Toronto and São Paulo at $4 a month each, Johannesburg and Tokyo at $8. Pro includes $4 of them. An added place checks every monitor on the account in its region's turn, and an ended one runs to the end of the period. A fixed IP is $4 a month on a place you added: that place's checks then leave from one IPv4 and one IPv6 address, shown on that page. Every account that added the place shares the address, so let it in together with a secret header on the monitor, never on its own. A monitor set to Check from a region where you added a place is checked from your places there alone, and not at all while none of them is online, since any other place would be turned away by its allowlist. Places and fixed IPs end when the plan no longer includes them.
Under the headline the page has a line per region with its places and what is down from it. Each strip has a lane per region under its bars, green where the region's checks got through and red where they did not, and response time by region folds out under it. Appearance has a switch that takes the regions off the public page.
Appearance also picks the regions whose readings count toward the page's response times: the medians, the bar heights, the slow threshold and slow alerts. Europe counts until you pick more, since a site answers each region at its own speed. The others still check up and down. A monitor can pick its own under Response times from, for a site served from several regions. Readings taken before a change keep counting as they did. While none of the chosen regions can check, Europe counts. Uptime counts picks the regions whose outages count against uptime; all of them until you leave one out.
Response times are medians. On the 24-hour range each bar is the median of the eleven readings around it in its region, and on the longer ranges the median of the hour or the day, so a single slow reading makes no spike and the bars hold still from one minute to the next. Above the strip are the median over the range and the figure 95% of readings came in under.
A region with more than one of our shared places reads at the level of the fastest of them. Each place is compared with the others over the hours both read, and the slower places' readings are scaled to the fastest's, so one place on a slower route to your site does not raise the numbers. The headline names that place: "from Aarhus". A place you bought is never scaled: its readings count as they are, and the headline names it too. With more than one region counted, each region is its own line and the bars and the median are their mean, so each region weighs the same whatever its number of checks. The API, status.json and the dashboard give the same 24-hour median, and every place still counts for up and down.
Vendor components
On Pro a component can follow a vendor's public status page. On the monitors tab, the add card has a tab for it: pick the vendor from the list, or paste the address of any status page hosted on Atlassian Statuspage, incident.io, Instatus, Better Stack, status.io, Sorry or StatOSS, or Slack's or Heroku's own. The page is read there and then, and you tick the parts of it you depend on, each shown with its state right now; every tick becomes one component. A large vendor's whole page is rarely all green, so tick parts where you can.
The vendor's page is read every five minutes. Operational shows as operational, degraded performance and partial outage as degraded, major outage as down. On the public page these components sit in a section of their own, Third-party services, unless you gave them a group, each with a line saying whose report it is. While the vendor has an incident open, its title takes that line, linked to the vendor, and the monitors tab offers to open an incident of your own from it. A vendor's trouble never moves your headline, your badge orstatus.json's page status: those speak for what you run. The headline adds one line naming the vendor instead. A vendor that cannot be read for 30 minutes, or that drops the part you followed, lets the component go back to operational, and you can set it by hand until the vendor answers again.
Groups
Monitors with the same group name are shown together on the page under a heading with the group's own state: all responding, or how many are slow or down. The subscribe form lists monitors with their group in front, so a group's members are easy to tick together.
What is kept
Every result is stored with its time, whether it passed, the response time, and the status code or the error. Each check also lands in an hourly total for its monitor, and hours older than 14 days are added up into days. Raw results are kept for 14 days (7 on Free); the totals are kept for 90 days, and a year on Pro, which is what the 7-day, 90-day and 1-year views read. The list of failed runs under a strip comes from the raw results, so on the longer views it covers the last 14 days. The plans page has the numbers side by side.