When One System Does Two Jobs: Verifying the Fast Path
At 11:36 AM on June 29, one click on Send Domain Instructions Email was taking the long way through our mail system.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
At 11:36 AM on June 29, one click on Send Domain Instructions Email was taking the long way through our mail system.
I kept saving goals to my agent's memory and they kept vanishing by the next session. The cause wasn't a flaky model. It was a silent truncation cap, and the fix is a pattern anyone running long-lived AI should steal.
Many apparent software failures are mismatches between the story a user reports, the state a system stores, and the assumptions a repair would make. A focused diagnostic can separate genuine drift from valid but unfamiliar data before anything is changed.
A single stuck proposal generated 17 cascading alerts and created a meta-deadlock. Here's how a zombie in the review queue taught us that every governance layer needs a fire exit.
At 2:54 p.m., the live production log showed the certificate auto-heal worker trying to process cf-heal-state.json as a command.
At 4:10 PM, a migrated site had the right DNS records and no HTTPS.
At 6:56 p.m., I had 12 unanswered design questions and a review tool I had built so another developer would not have to answer them in a ticket comment.
At 10 concurrent agents, the browser stopped being a tool and became a scheduling problem.
A bulk-edit feature passed its own code review on SQL injection, access control, and transactions. Sending it to an independent second reviewer found a second call site the fix never covered, a silent data-clearing bug, and a scope bug the first pass never asked about.
At 3:36 PM, we shipped a release because one organization-account tournament event was stuck and the exact trigger would not come back on demand.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.