---
title: "When Infrastructure Changes Faster Than DNS"
canonical: https://dxdev.com/blog/2026-08-09_dns-staleness-on-topology-migration/
datePublished: 2026-08-09
---
Yesterday I moved my vault knowledge server off a local tunnel and onto a real cloud endpoint. Same MCP tools, same agent, new address. I ran a full session against it. Vault search worked, vault reads worked, vault drops worked. I closed the laptop feeling like the migration was done.

This morning the same agent couldn't reach the same server. Not "returned an error." Refused to connect, then timed out, then connected to something that clearly wasn't my new box. I hadn't touched the config since the successful run twelve hours earlier.

## The instinct that wasted twenty minutes

My first move was to assume I'd broken the server itself. I checked the process, checked the port, checked the reverse proxy logs. Everything on the new box was healthy and waiting for traffic that never arrived. So the problem wasn't the destination. It was how my agent was finding the destination.

I ran a direct DNS lookup against the record my MCP config pointed at. It resolved to an IP that belonged to the old topology, the local tunnel endpoint I'd just retired. My own resolver was still handing out yesterday's answer. The record itself had been updated correctly at the source. What hadn't caught up was every cache sitting between my machine and the authoritative nameserver, each holding onto the old answer for whatever TTL it had cached before I made the change.

The deploy was done. DNS wasn't. Those are two different finish lines, and I'd only crossed one of them.

## I'd already read this exact failure mode

This is the part that annoys me most: I had the postmortem for this sitting in my own vault before it happened to me. During a large domain migration project I ran for customer sites, one client's domain went down specifically because a DNS flip landed before everything downstream was ready to receive it. That incident is why the client-managed domain rollout now runs on a self-serve model with a hard deadline instead of us flipping records on client-owned domains directly, and why there's a fail-closed guard so the migration engine can't take a client site offline if something's inconsistent at cutover.

I built guardrails against DNS staleness for other people's domains and then walked straight into the same problem on my own infrastructure, on a server I stood up in an afternoon with none of those guardrails, because it felt too small to need them.

## Why the lag is invisible until it isn't

The uncomfortable part isn't that DNS propagation takes time. Everyone knows that in the abstract. It's that the delay is completely invisible from where you're standing. Your deploy script exits zero. Your new server logs a clean startup. Every signal you actually watch says the migration succeeded. The only thing that hasn't caught up is a distributed cache you don't control and can't directly inspect, sitting between you and the agent that's about to try to use the thing you just built.

And it doesn't announce itself as staleness. It shows up as a connection failure, which looks exactly like a broken deploy, which sends you debugging the wrong layer.

This matters more for agent infrastructure than it did for the web servers I'm used to migrating, because the agent doesn't pause to wonder if something's off. It doesn't see "hey, this address feels wrong." It gets a timeout, retries, gets another timeout, and either fails the task or, worse, silently falls back to whatever it can still reach. If I'd kicked off an unattended overnight run against that endpoint instead of testing by hand first thing, I wouldn't have found a broken connection. I'd have found a run that quietly did nothing, or did the wrong thing, and I wouldn't have known until I went looking for output that wasn't there.

## What I actually do differently now

Any time I change network topology, local to cloud, direct to tunneled, one provider to another, I no longer trust that "the deploy succeeded" means "the path is live end to end." I check the record against the authoritative nameserver directly, not through my own resolver, because my resolver is exactly the thing most likely to be lying to me from cache. And I don't hand a freshly migrated endpoint to an agent for an unattended run until I've confirmed a clean connection from a machine that's never talked to the old address before, so there's no cache of its own to fool me.

None of this is exotic. It's the same lesson that large domain migration already taught me at the scale of a big batch of customer domains: cutover and propagation are two separate events, and the gap between them is exactly where things quietly break. I just hadn't internalized that the lesson applies at the scale of one server I stood up for myself, not just the scale of a migration with a project plan and a deadline attached to it.

The infrastructure changed in an afternoon. The DNS caught up on its own schedule. My job is to remember those are never the same clock.
