The Outage That Revealed a Two-Month Silent Failure
The status alert said the relay had been down for 22 minutes, and the relay was a box I thought I understood.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
The status alert said the relay had been down for 22 minutes, and the relay was a box I thought I understood.
A review or handoff is not complete when an agent says it saved the work. It is complete only when the receiving system can read the exact record that the next decision depends on.
We told a client their site would be live on the new relay by Friday. Friday came and went.
The invoice line item for GitHub Advanced Security did not match what the tool had done. The scan history for that repo said the last run was May 22.
A monitor built to catch our own outages paged a customer's dead third-party server as ours, six times in nine days, and I defended it before I was shown wrong.
Moving one league onto its sport's real code path surfaced a years-old tier gate that had simply never been built.
The first reply to a 375-person mailer arrived about three minutes after the send.
My live screen showed the same Discord conversation over and over. I went hunting for a bug that was writing duplicates. Nothing was: each card was a separate session, spawned five minutes apart.
A report built to catch the board and the code disagreeing before a release meeting had a blind spot exactly where seven real disagreements were hiding.
A daily alert flagged 17 tickets as inconsistent. Sixteen were false alarms, and the fix for them was about to run.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.