Fifty-five DMs before breakfast

Fifty-five Discord DMs were sitting in the channel by 7 AM, one bot message per line, all from the same script: the Cloudflare cert watcher that scans every domain on the platform and flags anything wrong with a certificate. My first read was that something had caught fire overnight. It hadn’t. Every one of the 55 was the same three customers who had never pointed their DNS at us, found again on that morning’s scan, same as the scan before it, and the one before that.

The finding kind was “off-edge”: Cloudflare has no record of ever serving that hostname, because the domain’s DNS still points somewhere else entirely. Real information, and zero shape for anyone at the platform to get woken up over. The customer has to fix their own DNS. There’s no relaunch to run, no fix commit to ship, nothing on our side to act on. Just a fact, restated by a cron job every scan cycle, forever, to whoever happens to be holding a phone.

Muting the finding instead of the shape

The fast fix, months earlier, had been to add off-edge to a small set called INFO_KINDS and skip paging on it. That worked, for off-edge specifically. It did nothing for stalled, a second finding kind the same watcher already emitted, for a domain that looked like it had never pointed DNS at us at all, not even a stale record. Same customer-side problem, same nothing-for-us-to-fix shape, different string. Nobody had connected the two, so stalled kept paging on its own schedule while off-edge sat quiet. I’d muted the finding I’d actually looked at, not the category it belonged to, and the second name for the same problem walked right past a filter written for the first one.

That’s the part that actually cost something: weeks of the exact noise the earlier fix was supposed to have ended, from a sibling of the thing already fixed, because the fix was keyed to a literal string instead of to what the string meant.

A ledger of already-told-you

The fix that actually held was a place to record that something had already been said, so nothing downstream had to re-derive it from scratch. A small module, three verbs: healed(key, what, venture) for a job that noticed a problem and fixed it itself, held(key, what, venture, reason) for a job that noticed a problem and chose not to page yet, resolved(key, venture) for closing one out. Every call appends one line to a ledger file. None of the three sends a DM on its own. Paging becomes a decision made once, upstream, by whoever decides whether a given finding is ours to act on.

off-edge and stalled both route through held() now, keyed by hostname, reason "customer DNS", and a morning digest reads the ledger back and shows them as a line item instead of a page. Three other finding kinds, expiring certs, stuck renewals, and a DNS TXT-drift check, got a second gate on top: they only page if the exact same signature also showed up on a scan from an earlier calendar day, not the same day. A cert mid-renewal that clears within the hour never crosses that gate. A second manual scan run an hour later doesn’t count as “survived a day” either. The gate reads dates, not timestamps. One that’s still stuck the next morning’s scan does cross it.

The same day surfaced an uglier version of the same problem, one layer down, in the cron wrapper every scheduled job calls. That wrapper already paged a failure the moment a job exited non-zero, keyed cron-fail:<name>, into an append-only log of one JSON line per run: name, return code, end timestamp. A separate health sweep, built earlier as a safety net for tasks that fail silently without the wrapper ever running, was independently polling the same Task Scheduler state and paging again under its own key, task-health:<venture>:<name>. Two different key shapes into the same alert gate meant the dedupe logic never recognized them as the same event, so one dead cron job produced two DMs. The fix was to make the sweep read the wrapper’s own history log before building its findings, and drop anything from its own list that already has a failed record there inside the last 24 hours.

The wrapper itself stopped paging on the first failure at all, for most jobs. It now reads the previous run’s own end record before writing this run’s, and pages only if that one also failed. A job that dies once and comes back clean next cycle gets logged as healed, not paged, because by the time a DM about it would land, it would already not be true. A handful of jobs, a daily backup, a weekly report where the “second occurrence” is a week away, opted back into paging on the first miss, because for those, waiting for a second failure is the wrong trade.

None of these three fixes look alike in the diff. All three answer the same question before sending anything: has a human already been told this, or would this alert be new information right now? Once something other than “assume no” is answering that question, most of what used to page turns out to have never needed a person at all.