How Autonomous Jobs Notify Humans
I found the failure during a production messaging-worker cutover, before I ever trusted SMS as the way an autonomous job talks back to a human.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
I found the failure during a production messaging-worker cutover, before I ever trusted SMS as the way an autonomous job talks back to a human.
355 chat sessions turned into 1,748 work units in one backfill, and the first thing the verifier caught was a git attribution label that was wrong.
On June 16, a 93-line script turned one publication rule into an executable gate.
The office lost access to its own Staff Center at 11:38 AM. Not a customer, not a hosting client. Us.
At 7:10 a.m., the apex relay health monitor fired. It fired again at 7:30. By 3:38 p.m., we had moved 25 brand domains, including the primary domain, off the old origin server and onto the apex rel...
On June 13, the RSS feed still had a draft leak, and the ready backlog contained 119 drafts.
The alert wasn't an alert. It was a customer telling us their site was down, at 3:57 PM on June 12th, four days after we'd already migrated their domain to Cloudflare and marked it done.
The vaultdrop tool disappeared from ChatGPT the moment I labeled it honestly.
I opened the ticket queue at 9:20 PM and closed it 87 hours and 21 minutes later, spread across the kind of overlapping sessions where you stop pretending you're doing one thing at a time.
A migrated domain sat in pending validation with an empty error field and a correctly published DNS record. I blamed the certificate authority. The token I was checking against had been rotated.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.