The bundle that was never there

At 6:33 AM, the daily sweep logged one line and moved on: vault drop failed, my home server unreachable over Tailscale, DNS resolution failing. The obvious read was that my PC was off. I filed it as a network blip and went back to the actual problem of the day, which was trying to get a second, freshly wiped machine talking to the rest of my setup.

That machine couldn’t restore. Not “restore was slow,” restore had nothing to restore from. And the reason took most of a day to find, because I spent the first chunk of it reviewing the wrong layer entirely.

The architecture that passed review

The backup design has four pieces. Everything, the full .secrets/ tree and every repo’s .env* file, lives in Cloudflare R2 via restic, versioned and deduplicated. A tiny cold-start bundle, just R2 access keys, the restic passphrase, and SSH keys, gets encrypted and dropped somewhere a fresh machine can reach without any of the rest existing yet. The passphrase that unlocks that bundle lives in a password manager. And the fresh box needs a short list of prereqs, restic and the ops-cli CLI on PATH, before any of it works.

That’s the design. It’s sound. I know it’s sound because I had two separate AI reviewers tear into it the same day the restore failed, and the review itself is where I burned the hours.

I’d floated my own alternative first: split the key material across two Google accounts, encrypted secrets in one, the unlock key in the other, so no single account compromise gets you both halves. I liked it enough to ask for a formal /multi-review before building it. Manus killed it in one paragraph: the split only defends against a single compromised account, and the normal bootstrap case is one Claude session that can read both accounts anyway, so the split adds operational fragility without adding real security. A password manager with a printed emergency copy does the same job with one less moving part. Codex’s independent pass caught something I’d have shipped without noticing: my runbook had rclone syncing to R2 when restic already speaks S3 directly, and R2 exposes a standard S3 endpoint. Adding rclone was one more install and one more failure surface sitting in the exact path that’s supposed to stay minimal. Both reviews were right. That session ran 8:37 PM to 1:19 PM the next day, sixteen hours and forty-two minutes, and it produced a genuinely better runbook.

It produced a better runbook for a system that was never actually deployed.

What the review couldn’t see

Because the real fault wasn’t in the design at all. When I finally searched my own Drive for the cold-start bundle, there was nothing there. No matching file, no .gpg, nothing. Somewhere between designing the four-layer scheme and calling it done, I’d run the review, written the doc, closed the ticket, and never actually executed the one manual step the whole thing depended on: generate the bundle, upload it, confirm it’s there. The fresh machine wasn’t failing to restore because of a bad architecture decision. It was failing because step two of a four-step plan had a checkbox nobody ever checked.

The fix, once I found the real problem, took minutes. ops-cli bootstrap bundle .tmp/cold-start.gpg --passphrase '<new passphrase>' on the healthy machine, upload the one resulting file to Drive, delete the local copy. On the fresh box, once restic and ops-cli were on PATH, ops-cli bootstrap unbundle pulls the keys and SSH material back down, then backup_r2.sh restore-all pulls everything else from R2. No architecture change. No split-key scheme. No rclone.

The two AI reviewers had done exactly what I asked them to do, and they’d done it well. Manus caught a real security tradeoff in my own idea before I built it. Codex caught a dependency I would have shipped. Neither of them could have caught “the artifact this whole design assumes exists doesn’t exist,” because nobody asked that question. I asked “is this design good,” and got a correct, detailed answer to that question. I never asked “did we actually do the thing,” and that’s a different question with a different kind of answer: not an opinion, a fact you check by looking.

A design review tells you whether a plan is sound if executed. It cannot tell you whether the plan was executed, because it has no way to see past the plan to the world. That check has to be separate, and it has to be dumb and literal: does the file exist at the path where the runbook says it should. I’ve since added that as its own automated check, run independent of any review, on a schedule, checking Drive directly rather than trusting that a closed ticket meant a completed action. The 6:33 AM failure message that started the day was a red herring about a network path. The real one, sitting unnoticed since the ticket closed, was a missing file that no amount of careful review was ever built to find.