The Safety Check I Trusted Wasn't One
Before running a rollback script against a live database, I switched on a syntax-only mode to make sure it was safe. It didn't check the syntax. It ran the script.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
Before running a rollback script against a live database, I switched on a syntax-only mode to make sure it was safe. It didn't check the syntax. It ran the script.
The roll-up said nothing was running while an executor was mid-turn. That was the first live test of the hub doorway, and the board was confidently wrong.
The fix took 1h 5m on a Thursday morning. Agents kept claiming content was on screen that I never saw.
The audit counted 621 reference files in docs/ and 57 files in dev-notes/, whose top level label said NOT for AI.
The load average sat at 4.2 on a box that should idle at 0.3. ps aux showed six next dev processes, oldest one three days alive, none of them attached to a terminal anyone was looking at.
We'd already built our internal ops-tracking system's Discord integration: daily, weekly, and monthly review cards, auto-posted on a schedule, pulled from the same session logs that drive the work ...
At 9:58 AM on July 22, our status board was still raising false-down pages when one monitoring signal disagreed with another.
At 9:58 AM on July 22, our status board was still raising false-down pages when one monitoring signal disagreed with another.
On July 22, I found an onboarding event that had fired on every page load since July 12. For 10 days, our signup funnel had been counting navigation as progress.
On July 22, every production mutation sent through our internal SQL tool's --write path rolled back when the connection closed. The command executed its SQL and returned without an exception.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.