Monitoring i status

An internal monitoring system with a public status page and incident handling. Uptime counted honestly - time without observation is "unknown", and maintenance windows do not inflate availability. Incidents live where the team already is, with a clear history for clients.

Monitoring i status
TL;DR

Monitoring that does not lie in its own favor. Time without observation is "unknown", not "up", and maintenance windows do not inflate availability. Add a public status page with an honest history and incident handling run where the team already is - on Discord, not in yet another tool.

Overview

Almost every status page you land on shows beautiful numbers. 100 percent uptime, all green, not a scratch. The trouble is that those numbers usually do not mean "everything worked", they mean "nobody was looking". Missing data gets counted as success, and the client gets a pretty picture with no backing in reality.

We built monitoring that does something harder than a pretty chart: it tells the truth, even when the truth is inconvenient. When we do not know how a service behaved, the result should read "I don't know", not "all is well". On top of that comes a public status page with an honest history and incident handling run where the team already sits. This case study is about counting availability without fooling yourself.

Green because nobody was looking

The mechanism of a typical status-page lie is trivial. If in a given period the probe collected no samples, and yet you count that period as healthy, your uptime will always look great. The less often you look, the better you come out, which is absurd: a tool for watching availability rewards you for not watching.

The effect is convenient and untrue. A green chart reassures the team and the client, even though underneath there could have been an hour nobody knows anything about. We wanted exactly the opposite stance: "I don't know" should mean "I don't know", not be quietly rounded up to "everything works".

A status page that is always green is not monitoring, it is wallpaper.

Three states instead of two

The whole honesty of this monitoring comes from one decision: a sample has three possible states, not two. Not just "up" and "down", but also "unknown" for the time when there simply was no observation. That distinction is the heart of everything else.

Only samples actually collected go into availability, that is the "up" and "down" ones. The "unknown" states are set aside, because you must not fake them in either direction. If there was no observation in a period, the result is not 100 percent and it is not zero. The result is empty, and that is exactly how the status page shows it.

availability.ts · ts
type Sample = 'up' | 'down' | 'unknown';

function availability(samples: Sample[]): number | null {
  const observed = samples.filter(s => s !== 'unknown');
  if (observed.length === 0) return null;
  const up = observed.filter(s => s === 'up').length;
  return up / observed.length;
}

Maintenance windows without stretching

A separate trap is scheduled downtime. It is tempting to count a maintenance window as healthy time, because "it was a planned break, not an outage". Except then the chart lies again, just more subtly: it pretends the service was running when it was deliberately gone.

We went with a middle that is honest in both directions. Maintenance windows are marked separately and do not count as downtime, because they are not an outage. But they also do not pretend everything was running then. The chart shows a gap marked as maintenance, not a fake green background. It is the same principle that governs the whole project: we do not round in our own direction.

Where 99.5 percent comes from (illustrative)

Available · 96%
Unavailable · 0%
Not observed · 4%
i
Note

On the pie chart above, availability is 995 out of 1000 observations, that is 99.5 percent, not 995 out of 1040. The forty unobserved samples go neither into the numerator nor the denominator, because we simply know nothing about that time. Counting them as available would inflate the result artificially, and as unavailable would deflate it just as artificially.

Incidents where the team already is

Monitoring that only counts is half the job. The other half starts the moment something actually breaks. Here the classic mistake is committed: incident handling is pushed into a separate tool nobody wants to log into in the middle of a fire. We went where the team already is, that is Discord.

The flow is short and needs no context switch. A service state change fires a notification. An incident is opened and updated straight from the channel, without logging into another panel. The public status page shows clients the timeline and the close, clearly and without dressing it up, as honestly as the uptime numbers.

1
Detection

a service state change fires a notification in the channel.

2
Driven from Discord

an incident is opened and updated from the place the team already sits in.

3
Public history

the status page shows clients the timeline and the close, without dressing it up.

What the finished product actually delivers

The number on the page means exactly what it says. When it reads 99.5 percent, that is 99.5 percent of actual observations, not an artifact of the probe sleeping for half a day. The client looks at the status page and sees an honest history: availability, downtime and maintenance windows kept apart, not eternal green that means nothing.

When something breaks, the team runs the incident from the place it already is, without logging into another tool at the worst possible moment. The whole thing holds to one principle that looks modest and is rare: monitoring you can trust, because it admits what it does not know instead of painting the gaps green.

Have a similar project?

Get in touch - a quote is free and comes back within an hour.