---
title: "Not Every Warning Deserves The Same Reaction"
canonical: https://dxdev.com/blog/2026-07-12_not-every-warning-deserves-the-same-reaction/
datePublished: 2026-07-12
---
A monitoring system paged the whole team because one single check came back a little slower than its threshold allowed. The very next check was completely normal. Nothing had actually gone down. The alert had turned one slightly slow moment into a full interruption.

The fix that mattered was not "make the alerts quieter." It was recognizing that a possible problem and a confirmed one are different kinds of evidence, and treating them as if they deserved the same response was the actual mistake.

## An alert can be technically correct and still be noise

The system had a line for "slow," and one single reading crossed it, so it reported a problem and paged. Nothing wrong with that from the system's point of view. It did exactly what it had been told to do.

From the point of view of the people who got interrupted, it was pure noise. Nothing was actually broken. There was one measurement outside a line, followed immediately by a normal one.

That distinction matters because a slow response can happen for all kinds of harmless reasons, a brief delay, a momentary bottleneck somewhere upstream, without the thing being monitored actually becoming unavailable. A system that pages on every single borderline reading teaches people to stop taking its warnings seriously. Once that happens, it has stopped doing the job it was built for.

## Two different kinds of bad news

The actual fix made the response asymmetric on purpose. A hard failure, something that doesn't respond at all, still triggers an immediate alert. That's deliberate. If something has genuinely stopped working, that is exactly the situation that needs the fastest possible response, and waiting for a second confirming check just to be extra sure would only add delay to the one failure mode that can least afford it.

A soft warning sign, like one slow but still-working response, is treated differently. It starts a streak instead of an immediate page. Only after it happens twice in a row does it become a real alert, and one normal reading in between resets that streak back to nothing. The single slow moment that caused the original interruption would, under the new rule, just be a measurement that came and went, never escalating into anything.

## Why the obvious shortcuts didn't work

The easiest response would have been to just make the "slow" threshold more forgiving and leave everything else the same. That helps a little, but a single unlucky reading above any threshold, however generous, can still trigger the same unnecessary interruption. It shifts the line. It doesn't fix the underlying problem.

Applying the same two-strikes rule to every kind of bad signal, including a hard failure, would have been consistent but wrong in a different way. Something that is actually down should not have to wait through a second check just because a separate, unrelated kind of alert had been annoying people lately.

A blanket rule that mutes or delays all alerts for a while after one goes off would have made things quieter too, at the cost of potentially hiding a real, separate problem that happened to show up shortly after an unrelated one already resolved itself.

## The rule

Ask what kind of evidence a given warning actually represents, and how much confirmation that specific kind of evidence deserves before it interrupts a person. A confirmed failure earns an immediate response. A single borderline reading earns a second look, not a page.
