---
title: "The 24-Hour Cache That Made Working Domains Look Broken"
canonical: https://dxdev.com/blog/2026-09-26_dns-verification-cache-staleness-trap/
datePublished: 2026-09-15
---
A customer's domain resolved correctly, served over HTTPS, and our staff Detail page still said "Awaiting DNS". It had said that for a day, and it would have said it for another one if nobody had looked.

## A cache that answered for 24 hours

The Detail page verifies setup by reading DNS. For client-managed domains, where the customer points their own records at us, the read went through a cache layer that could hand back an answer up to 24 hours old. So the sequence went like this:

1. The customer changes their records.
2. The cache still holds the pre-change answer.
3. The page reads the cache, sees records that don't match, and renders "Awaiting DNS".
4. The domain is already live, and support is looking at a status that says otherwise.

Nothing errored, and no log line said anything was wrong. The page was reporting a true fact about the cache and a false one about the domain.

## Stop asking the cache

We changed the page so that a client-managed domain is proven by a live DNS query at the moment the page loads. The cache no longer gets a vote on whether setup is done. If the live answer points at us, the domain is verified, whatever the cache last remembered.

It shipped as three hotfixes and a scan deploy. After the last one I checked two customer domains on prod and both flipped from "Awaiting DNS" to Secure without anyone touching them.

## The sweep that overcorrected

The wrong turn came in the sweep that walks the client-managed domains. It marked about 32 of them as gone when they were only at risk. Those are different states with different consequences. "At risk" means look at this soon. "Gone" means the domain no longer points at us, and it is the kind of flag someone acts on.

I caught it while verifying the earlier fixes, and it became a hotfix. I checked the corrected classification live on two domains before calling it done. Whether those flags persist in the database is still open on its own ticket, so for now I only know the page and the sweep agree, not that stored state has caught up.

The cost was small because nobody had acted on the flags yet. It still showed me that a fix which makes the display honest can make a nearby signal dishonest. The sweep and the page were reading the same world, and I had only corrected one of them.

## What the live check turned up

Once the page stopped trusting the cache, I ran a live sweep across every active domain to see what else it had been hiding. It found 30 customers still offline after the old origin was retired. Their records still pointed at an origin we had shut down. The stale cache and the stale page had been covering for that too, since a domain that looked "in progress" never looked broken.

Earlier the same day I had chased a different flavor of the same misdirection. Two custom-domain SSL failures looked like setup errors and were expired Cloudflare certificates. I reissued those and the other 72 hostnames stuck in the same limbo. In both cases the status on screen pointed at the customer's configuration when the fault was on our side of the line.

## A status that gates support needs a fresh read

A cache is fine for making a list page fast. It is the wrong source for a page that answers "is this customer live right now?", because that page is where a 24-hour-old answer becomes a day of someone telling a customer to check their DNS when their DNS was fine.
