Cron job monitoring
A URL your job requests when it finishes. Miss a ping and the job goes down on the page and an alert goes out, the same as any other outage.
Northwind is up.
The monitor is responding, 2 failed checks in the last 24 hours.
Nightly backup
Up for 6 h 51 min99.86% passed
1,438 checks, 2 failed checks
- Failed2 min, 2 checks
How a heartbeat works
A heartbeat monitor works the other way round. Instead of StatOSS requesting your service, your job requests StatOSS: the monitor's card shows a URL, and the job pings it at the end of every run. You set how often a ping is expected, say every hour for an hourly backup, and a grace period for runs that finish late. Once the period plus the grace pass without a ping, the next two evaluations in a row find it missing and the monitor goes down, the same two-in-a-row rule as every other check. The next ping brings it back at once. Before the first ping the clock runs from when the monitor was made.
This catches what a website monitor cannot: a cron entry that was deleted with the old server, a backup script that exits before uploading, a queue worker that hangs, a nightly report nobody notices is missing until the month end. Because it is a monitor like any other, it sits on the same status page as the site and the API, with an uptime figure and a strip of its own.
Sending the ping
Any request to the URL counts, with any method, from curl, wget, a fetch call, or a GitHub Actions step. Append it to the job so it runs only when the job succeeded:
0 3 * * * /usr/local/bin/backup.sh && curl -fsS https://statoss.com/api/heartbeat/<token>The && means a failed backup sends no ping, so the failure shows as a missed heartbeat. Heartbeats need no second opinion, since nothing is requested from our side; the monitor simply waits.
When it counts as down
A heartbeat monitor counts as down after two failed checks in a row and as back after one success. Before a failure counts, it is checked again in the same region; only a failure both checks see goes on the record. On Hobby and Pro every other region checks at the same moment, so the page can say down from North America instead of down. A single failed check still appears on the strip with its time and reason, it just does not change the state or send an alert.
Alerts
Each change of state sends one alert: down, and back, with how long it was out. Email is on every plan. Hobby and Pro add Slack, Discord, Microsoft Teams, Telegram, PagerDuty, Opsgenie, Pushover, ntfy and a signed webhook, plus a repeat notice every so many minutes while an outage lasts. Nothing is sent during a maintenance window.
What the status page shows
Every heartbeat monitor gets a strip on the public page. Each bar is one check, its height the response time; amber marks a timeout and red any other failure. Consecutive failures are grouped into one entry with the time and the reason, and the mark stays there for the whole range. A heartbeat monitor going down opens an incident on the page by itself and resolves it on recovery.