At 2:17 PM, I thought the SMS queue migration had been finished for several days.

The production cutover had passed a clean multi-day soak. The new queue was live. Nothing in the service behavior gave us a reason to keep the legacy path around. That is normally the moment a migration stops getting attention. The new system works, the dashboard is quiet, and the team moves on.

Then I started the retirement pass.

The queue was only one part of what we changed. The final production state touched legacy services, a CI runner, working folders, an API repository, the main application repository, a runtime registry in our shared vault, and a standalone repository. The deployment had succeeded. The map of the deployment had not caught up.

That was the actual last mile. Not another deploy, but an audit of every place a future engineer or automation could learn the wrong version of reality.

I would have told anyone who asked, in those several days, that the migration was done. I believed it. The production traffic supported that belief. Nothing about a clean soak told me that three different repositories and a shared registry still disagreed with each other and with the runtime, and I only found out because I went looking for a reason other than doubt: closing out the ticket properly, not chasing a suspicion.

The cutover was not the finish line

We did not shut down the old path the first time the new queue worked. We waited for the multi-day production soak, verified the cutover was clean, and only then started removing things. That sequence mattered. A queue migration can look healthy during a short test and still fail when a retry path, a scheduled job, or a quieter production condition finally appears.

Once the soak passed, we decommissioned the legacy services, the CI runner, and the folders that existed only to support the old path. Each removal closed a different kind of gap. Leaving the runner alive would preserve a way to build or run stale code. Leaving the folders around would make the retired deployment look maintained. Leaving services registered would tell an operator that they were still available to start.

The tempting alternative was to call the migration done after the soak and file the cleanup for later. We did not take it. Stale infrastructure begins teaching people immediately. A service name in a runtime registry is not harmless documentation. It is an operational instruction. An idle CI runner is not harmless capacity. It is a future diagnostic branch, one more thing someone has to account for under pressure.

We treated decommissioning as part of the production change, not housekeeping after it.

Three repositories, three competing stories

The documentation audit had to cross three locations because the system was described in three locations. The API repository carried the code and service-facing integration around the queue. The main application repository carried the application-level operational picture. The shared vault runtime registry answered the most important question, what actually exists and can run.

The final WinSW service state was the anchor. We started there, then compared each repository and the registry against it. We did not rewrite prose until it sounded current. We asked whether every record described the same runtime that had survived production.

That distinction changes the audit. If you start from documents, you ask whether each page seems plausible. If you start from runtime state, you ask a stricter question: can this line make someone look for, start, deploy, or depend on something that no longer exists?

That is how stale service names surfaced. They were not necessarily breaking production in that moment. They were believable, which made them more dangerous. An engineer following one later could lose time on an abandoned path. An automation reading an old reference could preserve the wrong assumption. The production service might be correct while the organization is still carrying an expired model of it.

Two orphaned services, two different problems

The audit also uncovered two services that did not belong in the final picture.

One was a dead proof of concept. It was harmless in the narrow sense that it was not doing live work and did not belong to the queue. But that did not justify leaving it documented as runtime. A dead proof of concept makes the system look more complicated than it is, and it turns an old experiment into a false maintenance obligation.

The other was future scaffolding that had never been built. This was a different mistake. It was not an obsolete production component. It was an idea that had acquired the appearance of a deployed component because it entered the infrastructure record too early.

The cleanup was the same in one important way. Neither service could be presented as live runtime. The audit was not just searching for old names. It was checking the lifecycle status of every entry against reality.

I now use four explicit categories when I find a service entry during a migration audit: live, retired, experimental, or future. A legacy service replaced by the cutover is retired and its operational references need removal or correction. A proof of concept is experimental and cannot sit in live runtime documentation without that label. Future scaffolding stays future until a real deployment exists. The distinction is simple, but a runtime registry that mixes those categories is not a registry. It is a notebook of possibilities.

Archive the history, remove the ambiguity

The standalone repository was the clearest artifact of the old architecture. Once the new queue had survived its soak and the legacy services, runner, and folders were gone, that repository no longer owned an active deployment path. We archived it rather than leaving it in the normal working set.

That is different from ignoring it. Archiving says the code is retained as history, but it is not a current operational answer. This matters when someone searches across repositories during an incident, or when automation enumerates projects and finds old code with a familiar service name.

Keeping the repository active would have preserved another competing source of truth. Deleting it outright would have thrown away useful history. Archiving separated those concerns.

The closing pass I want on every migration

The migration was not complete when the new SMS queue processed production traffic. It was complete when the old runtime was removed, the CI runner and folders tied to it were gone, the standalone repository was archived, and the API repository, main application repository, and shared runtime registry all described the same final WinSW state.

My closing pass is now concrete. Let the new path survive a real production soak. Decommission every retired service and the tooling that exists only for it. Use final runtime state as the audit baseline, then compare every repository and registry against that baseline. Classify every stray entry as retired, experimental, future, or live. Archive code that remains useful as history but no longer owns an active deployment path.

A clean deployment proves that the new thing works. This audit is what proved the old thing was actually gone: seven places checked against one final WinSW state, two orphaned services relabeled out of live runtime, one standalone repository archived rather than left to compete with the truth.