---
title: "Asymmetric Alerting: Stop Paging on Transient Blips"
canonical: https://dxdev.com/blog/2026-08-22_transient-alert-flap/
datePublished: 2026-07-12
---
At 4:59 PM, a Cloudflare edge alert marked a site degraded because one probe came back **53 ms** over the slow threshold. The site was up the whole time. The next cycle was clean. Our alert logic had turned a one-cycle latency blip into a page.

I fixed it in 17 minutes, but the fix was not just “make alerts quieter.” I changed the status board so it treats a possible degradation and a confirmed outage as different events.

A slow response is evidence that something may be wrong. A failed availability check is evidence that something is wrong right now. Those states deserve different confirmation rules.

## The page was technically accurate and operationally wrong

The board had a slow threshold and a degraded state. One probe exceeded that threshold, so the board reported degradation and paged. From the code’s point of view, it did exactly what it was told to do.

From an operator’s point of view, the alert was noise. There was no customer-visible outage. The site did not go down. There was one measurement outside a line, followed by a normal measurement.

That distinction matters because latency is noisy in a way a hard availability failure is not. A probe can hit a slow edge, contend for a connection, or absorb a short network stall. The response can be late without the application becoming unavailable. A status board that pages on every isolated latency violation trains people to discount its degraded state. Once that happens, the board loses the job we built it to do.

The diagnostic path was short because the symptom was specific. I checked whether the site had actually gone down. It had not. I checked the alert event. One probe was **53 ms** over the old slow threshold. I checked the following cycle. It recovered. That made the incident a one-cycle flap, not a sustained performance problem.

The original threshold was also too tight for the meaning we wanted “slow” to carry. I raised it to **2500 ms**. That change reduces sensitivity, but it does not solve the alerting problem by itself. A single response over 2500 ms can still be a transient. The real change was in what a degraded sample is allowed to do.

## Two bad cycles for degraded, one for down

The new policy is intentionally asymmetric:

```yaml
slow_threshold_ms: 2500
page_on:
  down: first_failed_cycle
  degraded: second_consecutive_bad_cycle
```

A failed availability check still pages immediately. That is fail-closed behavior. If the board cannot reach a service, I want the first failed cycle to create an incident. Waiting for a second check would buy a little protection from probe noise while adding delay to the failure mode that needs the fastest response.

A slow response now starts a degraded streak instead of creating a page. A second consecutive slow cycle confirms it. A clean cycle resets the streak. The single 53 ms excursion that caused the 4:59 PM page would now appear as a measurement, then disappear without escalating the whole team.

In plain control flow, the policy is this:

```text
if availability_check_failed:
    page_down_immediately
else if response_time_ms > 2500:
    increment_consecutive_degraded_cycles
    if consecutive_degraded_cycles >= 2:
        page_degraded
else:
    reset_consecutive_degraded_cycles
```

The ordering is part of the design. The `down` path is checked first and has no debounce. The `degraded` path is explicitly stateful. Treating both conditions as one generic “unhealthy” state would have erased the difference that matters.

## The fixes I did not choose

The most obvious response was to raise the slow threshold and leave the rest alone. I did raise it to 2500 ms, but threshold tuning alone would still page on an isolated slow sample. It changes where the line sits. It does not distinguish a momentary measurement from a sustained condition.

I also could have applied the same two-cycle rule to everything. That would have made the alert logic superficially consistent, but operationally worse. A service that is actually down should not wait through another monitoring interval because one transient latency event was annoying.

A global alert cooldown would have been another way to reduce noise. It would also have suppressed information indiscriminately. A real outage immediately after a benign flap could be hidden behind the same timer. Cooldowns control the delivery channel. Consecutive-cycle confirmation controls the evidence required for a specific state. I wanted the latter.

The asymmetric rule is slightly more complicated than a single threshold, but it is legible. Anyone reading it can answer the important question: how much evidence does this state require before it wakes someone up?

## Verify the behavior, not just the branch

After changing the threshold and debounce logic, I verified the new behavior live. The point was not simply that the status board still rendered. The policy had to preserve both sides of the tradeoff: one bad latency cycle should not page, two consecutive bad cycles should, and a down condition should still page on the first failure.

That is the general pattern I am taking from this one alert. Do not ask whether an alert is “sensitive enough” in the abstract. Ask what kind of evidence each failure state produces, how often that evidence flaps, and what delay is acceptable before someone acts.

A slow response can be a warning. A failed check is an incident. The alerting logic should know the difference.
