A server needed a restart it had put off for weeks. Two months of updates had been staged on it, and staged updates do nothing until the machine restarts. I wrote down how long the website would be down, and the number I wrote was a guess.

Two guesses in a row

An early version of the plan said about three minutes. It read like a fact because that is how downtime windows normally read. When I checked what was actually waiting on the machine, that looked too small for two months of updates, so the plan was changed to a wide range of 20 to 45 minutes, with a note about what happens if someone trusts the old number: they stand there at minute 20, unsure whether the machine has died.

Neither number had been measured. One was too small and the other turned out too large.

What the run showed

The restart happened early in the morning. Someone had to press the button, and that was me. The watching was done by an AI agent, which stood outside the machine and recorded what the site did. Partway through, that agent session kept failing on errors from the AI service, and a second one took over the watch.

The site was down for about two minutes. Then it went down again for about two more, without anyone doing anything. About four minutes of outage in total, across two restarts, and nothing like 20 to 45.

The second restart was not in any version of the plan. The machine’s own update installer had queued a follow-up restart to finish the job, and that is normal behaviour once you know to look for it. Nobody had looked, because nobody had run this machine through this procedure before. Whoever was watching would have taken the second drop for a failure.

What went into the plan afterwards

The same morning the written procedure got four corrections:

  • A warning that the machine restarts twice, how to tell the second restart from a real failure, and a rule not to run the after-restart checks between the two.
  • The wider range, kept, with the one real measurement added as a data point instead of replacing the estimate.
  • A note that the machine’s memory ceiling had changed by itself across the restart, from 40.91 GB to 36.66 GB, so a before and after comparison should expect it.
  • Three warnings that appear on every start-up of that machine, before and after, listed as known noise so nobody chases them.

I also added a check that the updates had actually installed. A cleared flag does not prove they landed.

Where AI fits

The AI agent could watch, time and write down what it saw. It could compare the run to the plan and list where the two disagreed. Someone still chose the day, pressed the button and decided whether the plan was fit to rely on.

The human decision

Whether to run the risky step, and when, stayed with me. So did deciding what a guess is allowed to look like in a plan. Judging that “about three minutes” was not good enough to act on took a person reading the situation, not a tool.

The lesson

Mark every number in a plan as measured or guessed, and say which run it came from. The runbook now carries one measured figure, about two minutes per restart, recorded as a data point from a single run next to the wide estimate. That is as much as the next person is entitled to trust.

The Build Log companion covers the restart mechanics and the corrections to the procedure.