---
title: "The Recovery Job That Would Have Sent 3.5 Million Messages to the Database, One at a Time"
canonical: https://dxdev.com/blog/2026-01-27_thirty-minutes-and-3-5-million-messages/
datePublished: 2026-01-27
---
Thirty minutes was all I had before the recovery job would be stopped, and the first version was headed toward 3.5 million separate messages to the database.

A sports site had been deleted. Bringing it back meant copying its old records from a backup: teams, games, scores, and the other pieces that make a season hang together. The copy itself was only half the problem. When a game was put back, it received a new number. Every score that belonged to that game had to be pointed at the new number, not the old one. Otherwise the site could look repaired while its results were tied to the wrong game, or to no game at all.

I had already built the careful version. For every old game number and new game number, it sent a separate instruction to change one kind of record. It was easy to follow. It was also the kind of approach that feels sensible when you are handling one team at a time. You can picture each change, check it, and move to the next one.

Then the real job was a league with 327 teams.

The same small instruction had to be repeated for many games, many kinds of records, and the saved copies of older records too. What had been a tidy little loop became about 3.5 million individual requests. Each one had to leave the recovery job, reach the database, get an answer, and make the trip back before the next one could start. None of those pauses looked dramatic by itself. Put millions of them in a row, though, and the job could not finish inside its thirty minute limit.

That was the cost of my first answer. It was not wrong. It was too slow to be useful when the number of teams got big.

The fix was a temporary table. That is just a short-lived list inside the database, used while a job is running and then thrown away. I put every old game number beside its new game number in that list, along with the team it belonged to.

That changed the conversation completely. Instead of asking the database to make one change, then another, then another, I gave it the whole address book first. The list was loaded in groups of 500 entries so the instructions stayed manageable. Then each kind of record received one broad instruction: find every item whose team and old game number match this list, and replace that old number with the new one next to it.

It is the difference between changing the forwarding address on one envelope at a time and handing the post office a clean sheet of every address that changed. The destination still matters. The names still matter. But the person doing the work is no longer stopping to make a new trip for every envelope.

There was one part I would not trade away for speed. Different teams can reuse the same old game number. A change based only on the number could quietly point one team’s score at another team’s game. So every match used two pieces of information together: the team name and the old game number. The faster version kept that same rule. It changed how the instructions were delivered, not what counted as a safe match.

I also did not force the bigger method onto every recovery. A single account does not need the extra setup of making and filling a short-lived list. For that smaller case, the original approach stays easier to trace. A plan-only run that makes no changes also stays on the no-write path. The new route only turns on when there are at least two accounts and the recovery is actually allowed to change records.

I made the job say which route it had taken and report its progress as accounts were handled. That sounds small, but it matters when a recovery takes long enough for somebody to wonder whether it is working or stuck. A plain line saying that the faster route is in use gives the next person something real to check.

One more problem was hiding behind the faster instructions. Giving a database a long list does not guarantee that it can find the matching records quickly. It needs an index, which is a lookup card that helps it find the right drawer without opening every drawer in the room. In this case, the useful card paired the team name with the game number. Without that, a few large instructions could still waste time searching through whole tables. The list and the lookup card had to work together.

The result was not a cleverer loop. It was a different shape of work. The recovery went from roughly 3.5 million separate messages to a few hundred: a couple of short-lived lists, several groups of 500 entries, and one broad change for each kind of record. The smaller jobs kept their readable path. The large league had a path that could actually finish.

The real fix was giving the database one list instead of millions of separate requests: a temporary table of every old-to-new game number, batched 500 at a time, joined on team and old id together so no team's scores could land on another team's game. That's the number worth keeping: 3.5 million round trips became a few hundred, not because the loop got faster, but because it stopped being a loop.
