When Operations Can Fail Partway: Use Rollback, Not Error Handling
The reconcile only ever retries removed rows. That's the whole bug in one sentence, and it took a half-written domain to find it.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
The reconcile only ever retries removed rows. That's the whole bug in one sentence, and it took a half-written domain to find it.
The alert never fired. What fired instead was a teammate asking twice, five days apart, whether his preview build had actually gone out.
The audit was clean: 211 wrapper merges checked, zero missing a tag on origin. Ninety days of history, nothing untagged.
The upload log for August 29th shows a phone photo landing at 1.43 MB, a JPEG the customer picked at 438 KB on disk.
The deploy had been dead all night. Not a code error, not a failed test: a GitHub Actions billing block, silent from inside the repo and only visible on GitHub's org settings page.
The manus-watch commit says it missed 152 real credits. The commit before it says a shared alert key ate a real 28-credit charge.
The alert log at 12:49 PM said it plainly: three separate things sat behind the agent dock feeling broken and mysterious. That framing only exists because the debugging didn't start that way.
The alert said the failure was a cache issue. It wasn't. The job had run zero steps, in two seconds, twice in a row, and the actual reason was sitting in a GitHub check-run annotation our triage sc...
A text-only grep and a text-only scanner both passed while a real name sat in the EXIF data of 13 images on a pseudonymous site. An independent check found it.
I put my own card through the live payment account, the first payment it had ever taken.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.