Uptime monitoring
Five minute checks, confirmed failures, and an incident log.
Uptime monitoring asks one question repeatedly: is the site answering. It is separate from analysis, because the best-optimised page in the world scores nothing while it is returning a 502. Available from the Pro plan.
Its counterpart is URL monitoring, which re-runs the full analysis of a page on a timetable. Uptime watches whether the site is there; URL monitoring watches what is on it.
How it works
| Every five minutes | Each monitored site is requested on a five minute cycle, around the clock. |
| A real HTTP request | Not a ping. We make the request a browser would and read the status code that comes back, so a server that is running but returning errors is reported as down, because to a visitor it is. |
| Response time recorded every check | Kept alongside the status, so you can see a site getting slower before it becomes a site that is off. |
Why a failure is confirmed before it is reported
A single failed request is not an outage. Networks drop packets, servers restart, a deploy takes four seconds. We check again before saying anything, and only a second consecutive failure opens an incident.
This is the difference between a monitor you act on and one you learn to ignore. A monitor that alerts on every dropped packet trains you to dismiss it, and then it is worth nothing on the day it is right.
What a check can return
| Status | Means |
|---|---|
| Up | The site answered with a success status. |
| Down | A server error, a connection refused, a timeout, or a DNS failure. Confirmed by a second check. |
| Slow | The site answered, but took long enough that visitors would notice. Recorded, and not counted as an outage. |
| Paused | Checking is suspended by you. Not counted either way. |
Inside one monitor
The list gives you a status and a strip of bars. Opening a site is where the detail is.
A date range across the top drives everything below it: 24 hours, 7, 30 or 90 days, or two dates of your own. The four figures restate themselves for whatever window you pick.
| Figure | Means |
|---|---|
| Uptime | The share of checks in the window that succeeded, with the number of checks behind it so you can see how much the percentage is worth. |
| Downtime | How long the site was down in total, and across how many separate outages. Five minutes in one outage and five minutes in ten are very different problems. |
| Average response | The mean of the successful checks, with the fastest and slowest beside it. The range matters more than the mean: a steady 600ms and a mean of 600ms that swings between 200 and 5,000 are not the same site. |
| Failed checks | Individual checks that failed, counted before confirmation. Always at least as high as the outage count, and usually higher. |
Failed checks being higher than outages is the confirmation rule doing its job. Each outage needs two consecutive failures, so single failures are counted here and never raised an alert. A large gap between the two numbers means the connection is flaky rather than the site being down.
Every check, and response time
The bar strip is one mark per check across the window, green for a success and red for a failure, so a pattern is visible as a shape before you have read a number. Beneath it the response time is plotted with its typical value and range called out.
Read them together. A run of slower responses ending in red bars is a server running out of something; red bars with no slowdown in front of them is usually a deploy or a network event.
Incidents
An incident opens on a confirmed failure and closes on the next success. It records when it started, when it ended, how long it lasted and what the site was returning throughout. You can add a note to one, which is how the log stays useful a month later when nobody remembers what the cause was.
Reading the history
Availability is shown over the last 24 hours, 7, 30 or 90 days, with response time alongside it. The two together are worth more than either alone.
| Shape | Usually |
|---|---|
| Brief failures at the same time daily | A backup, a cron job, or a log rotation competing with the web server. |
| Response time climbing over weeks | A database growing without an index, or a cache that is no longer big enough. |
| Failures clustered around deploys | No zero-downtime deployment. Small window, entirely avoidable. |
| One long outage | Worth a note on the incident while you still remember what it was. |
Pausing
Pause a monitor before planned maintenance so the window does not count against your availability figure or raise an alert for something you are doing on purpose. A paused monitor is not checked and is clearly marked as paused rather than as healthy.
Pro monitors 3 sites, Agency monitors 10. A site is a hostname, so monitoring onbixo.com covers that host rather than each page on it. Everything here is available over the API: see Website uptime.