At 9:58 AM on July 22, our status board was still raising false-down pages when one monitoring signal disagreed with another.

The service was not actually down. The failure was in how we turned evidence into an alarm. A collector reported a problem, the live web probe did not confirm it, and the board treated the collector’s report as enough to announce an outage.

By 1:09 PM, three hours and 11 minutes after I started, the board was running a different rule in production: it only alarms when the collector and the live probe agree that the service is down. We also gave it a queryable event log, so the next disagreement has a history instead of a mystery.

The alert had one sensor too many

The collector and the live probe were not redundant copies of the same check. They answered related questions from different places in the stack.

The collector supplied one view of service health. The live web probe tested what a user could reach on the web. On the bad pages, those views split. One said down. One did not. The board collapsed that split into a single down state.

That was the diagnostic path. I did not need a more elaborate incident narrative to explain it. The false-down symptom lined up with a simple fact: our alert decision did not preserve disagreement. It converted one negative signal into a page before the second signal had any say.

A monitoring system can be correct about an individual observation and still wrong to escalate from it. A collector can surface a transient failure, a stale result, or a problem limited to its own path. A live probe can remain healthy through all of that. When the system reports those two observations as one unqualified outage, it hides the only fact that mattered: the sensors disagreed.

Make the decision boundary explicit

The production change was small, but it changed the meaning of the alarm. In effect, the rule is now:

alarm = collector_down AND live_probe_down

The collector still matters. The live probe still matters. Neither one can page by itself.

That left three states worth distinguishing. If both report down, the board alarms. If both report healthy, it stays quiet. If one reports down and the other reports healthy, it records the event without claiming there is an outage.

The alternatives were straightforward, and neither fixed the actual defect. We could have kept the collector as the single source of truth. That preserves a fast alarm, but it also preserves the false-down path that was already paging us. We could have made the live web probe the only authority. That replaces one blind spot with another and throws away useful collector evidence. We could have softened the threshold around one metric, but that just changes how much noise a single sensor can create.

Consensus was the narrower change. It did not pretend either signal was perfect. It asked for corroboration before crossing the line from observation to alarm.

The event log is part of the fix

Suppressing a false alarm is not enough if the evidence disappears with it. The mismatch is still operational data. It tells us that the collector and the live probe did not see the same thing, and it gives us a starting point if the disagreement becomes a pattern.

The board now keeps those events in a queryable log for troubleshooting history. That matters because the next investigation does not begin with a vague report that the alerting system was noisy. We can ask what each signal said at the time, find the disagreement, and decide whether the issue belongs in the collector, the probe, or the decision rule.

We deployed the change and verified it in production. The constant false-alarm path went to zero, not because we ignored a bad signal, but because we stopped letting one sensor speak for the entire system.

The useful change was not adding another dashboard or a more complicated severity model. It was refusing to call a disagreement an outage.