---
title: "The Batch That Was Aging Out With No Alarm Attached"
canonical: https://dxdev.com/blog/2026-08-06_the-batch-aging-out-with-no-alarm/
datePublished: 2026-08-06
---
Nearly a thousand of our TLS certificates for customer-facing sites were sitting on a validation method that needed a fresh DNS record from the customer's own domain every 90 days, and nothing in place would have noticed a specific one about to fail before the renewal actually did.

## How it happened

Months earlier, a one-time bulk migration onto a CDN's managed-certificate product had registered close to a thousand hostnames using the older of two validation methods, the kind that checks a special record sitting in the customer's own DNS. Every hostname added since had gone in on the newer method by default, the kind the CDN answers itself at its own edge, with no DNS record required from anyone. Nobody had chosen the older method on purpose for that batch. It was just what the migration tooling happened to write, and it sat there working quietly for months, because a certificate on that method only actually needs the DNS record present at the moment it renews.

## Why the existing safety net wouldn't have caught it

We already had a script that watches for certificates stuck mid-renewal and repairs them automatically. It only acts on a hostname it can already see failing, one sitting in a stuck or timed-out state. A hostname on the older method that hasn't hit its 90-day mark yet doesn't look stuck to anything. It looks exactly like every healthy certificate, right up until the DNS record it needs isn't there anymore and the renewal actually fails.

## What made it visible

The CDN itself surfaced it, sending "action required" notices as each 90-day window approached. Reading a handful of those closely turned up the real shape of it: a batch, not a one-off, all sharing the same migration origin, all coming due around the same season.

## The fix, and the number that mattered

Every hostname on the old method got flipped to the new one in a single pass. Zero of the flips failed. The estate afterward was clean, all of it on the method that needs nothing from any customer, ever, at renewal.

## The gap that made this possible in the first place

The reactive script and this fix solve different problems, and neither replaces the other. One repairs a renewal that has already broken. Nothing before this existed to notice a batch quietly approaching that break before any of them actually hit it. That's what got built alongside the fix: a watchdog that reads the whole estate on its own schedule and says something the moment a certificate is close to expiring without a completed renewal, a hostname has drifted back onto the old method, or an onboarding never got a certificate at all, instead of waiting for a vendor's warning email or a customer to notice first.

## The lesson underneath it

A fix that repairs failures after they happen and a check that notices trouble before it happens are not substitutes for each other, and a system can look healthy for a long time running on only the first one. The certificates that had renewed fine every 90 days for months looked exactly like the ones that never would have, right up until the day one of them didn't.
