On May 20, I had five separate documents explaining how work phases were supposed to operate in my agent system. None of them agreed on every detail. The agents had been reading whichever one appeared in context and filling the gaps with their own interpretation.
That is the bug. Not missing features, not wrong prompts. Five near-duplicate explanations of the same workflow, none of them canonical, each one producing slightly different agent behavior depending on which session loaded which file.
I had been treating workflow as a thing I could re-explain whenever it came up. That turned out to be a slow form of decay.
the canon problem
My AI ops system handles work in phases. A task opens, gets triaged, gets estimated, gets worked, gets reviewed, gets closed. Each phase has criteria, transitions, JIRA field states, and rules about what can and cannot happen at that moment. This is not complicated in concept. But the rules had accumulated across session prompts, memory entries, CLAUDE.md sections, and draft architecture docs that never got promoted anywhere.
When a session started, the agent would pull whatever was in context and synthesize. Most of the time the synthesis was close enough. Occasionally it was not. A field that should have been set during the coding phase got set at close. A transition fired before the work was verified. A reviewer got skipped because the session “remembered” a shortcut from a different task context.
The signal came from the correction patterns. My weekly retro script walks recent transcripts and surfaces recurring mismatches. The same family of errors kept showing up: agents inventing phase rules that had never been agreed on.
The root cause was not that the agents were bad reasoners. It was that they had nothing authoritative to reason from.
writing phase_contracts.md until it was promotable
I started a document called Phase_Contracts.md. The first draft took about 40 minutes and was roughly right. Then I did something I had been putting off: I ran it through a full multi-reviewer cascade before treating it as canonical.
The sequence was Codex first, then Manus. Codex is synchronous and fast, good at catching logical gaps and undefined states. Manus runs async and sees Codex’s pass before it starts, which means it tends to catch the things Codex normalized rather than challenged. Three rounds of back-and-forth across both reviewers. The document went through five versions.
Version 1 had the phase names right but left transition criteria ambiguous. Version 3 caught that I had been using “ready for review” to mean two different things in two different contexts. Version 5 had an Appendix C.
Appendix C was not planned. It appeared because the reviewers kept generating follow-up actions that did not belong inside the contract itself. By version 5 there were 11 of them: specific code changes, field validations, automation hooks that the contract implied but that nothing had actually implemented. A real contract generates implementation pressure. A draft that stays a draft avoids it.
Promoting the document into the vault meant committing to resolving those 11 items, or at minimum acknowledging them as open debt. That is the moment a workflow document stops being aspirational.
the jira split that should have happened earlier
One of the five version rounds surfaced something I had been glossing over. I run two JIRA instances. HQ is Atlassian Cloud, my personal project space. ITEM is a self-hosted JIRA Server instance running for the sports-league-management product I own. They are different systems with different API shapes, different custom field IDs for the same concepts, and different transition IDs for the same state changes.
My agent code had been treating them as nearly the same and patching the differences inline. That works until it does not. The failure mode is silent data corruption: a field write that succeeds on one instance silently no-ops or corrupts on the other, and the agent reports success because the API returned 200.
The phase contract forced me to be explicit. customfield_10511 is Dev Notes on the platform server. The same concept on HQ Cloud lives at a different field ID. Transition 101 means CODING to STAGING/LIVE on the platform. On HQ the same logical state change uses a different numeric ID. These are not details I can normalize away. They are structural differences that need to live in a config layer, not in ad-hoc patches spread across session prompts.
I split the field and transition maps by instance. Each JIRA target now has its own config block. The agents read the right block for the right instance, and the contract names which block applies at which phase.
This felt like plumbing. Silent data corruption because two instances share a field map is not plumbing. It is a correctness bug dressed up as a config detail.
turning the review flow into a skill
The Codex-then-Manus cascade that produced the five versions of Phase_Contracts.md is not a novel idea. I had done it before on other documents. But I had done it manually each time: open session, paste document, run reviewer, copy output, open new session, repeat. The sequence was not written down. It was not rerunnable. It was a thing I did, not a thing I had.
So I turned it into a skill: /multi-review. The skill takes a document path or text block, dispatches Codex first, waits for output, then dispatches Manus with the Codex pass prepended as context. It saves each reviewer’s output to .runs/multi-review/ with a timestamp and version tag. At the end it summarizes what changed and what remains open.
The reason to do this is not efficiency in the moment. The Codex-Manus sequence on a 1,500-word document takes about 20 minutes regardless. The reason is that a process you cannot rerun is not a process. It is a story you tell about something that happened once.
The next time I want to promote a document, I run /multi-review against the draft. The sequence is deterministic. The output is auditable. If a future session asks why a particular field rule works a certain way, I can point at .runs/multi-review/ and show the version history of the decision.
what the 11 appendix c items actually mean
The contract’s Appendix C is not a backlog. It is a gap register. The difference is that a backlog is a pile of things I might do. A gap register is a list of places where the system is currently operating on trust instead of enforcement.
Item 2: field validation at transition time, check that Dev Notes is populated before the CODING transition fires, not after. Currently the agent is supposed to do this by convention. Convention is not enforcement.
Item 7: phase duration logging, record when each phase starts and ends per ticket, so the retro has actual data instead of transcript inference.
Item 11: Appendix C itself should be a living section, updated when items close and when new gaps surface.
I have not resolved all 11 items. The contract is not in fully production-trustworthy state just because the document is clean. But the gaps have names. They are tracked. That is the actual value of writing the contract down: not the document, but the gap list the document generates when you take it seriously.
what changed after promotion
The week after I promoted Phase_Contracts.md to the vault, I ran three work sessions on active tickets. In all three, when the agent reached a phase transition, it cited the contract directly. “Per Phase_Contracts.md section 3.2, this transition requires a verified Dev Notes entry.” It did not invent a rule. It read one.
Two of those sessions had a moment where the agent noted a discrepancy between the contract and an older memory entry. Both times it flagged the conflict and asked which was authoritative. That is the right behavior. The contract won both times.
The third session caught a bug in the contract itself. Section 4’s criteria for the review phase used “ready for review” without defining what “ready” meant, the same ambiguity that version 3 of the draft had tried to fix and had not fully resolved. I updated the contract and ran /multi-review again to verify the fix held.
The system is not done. But it is no longer improvising.
why “if you explain it every time, it is not a workflow yet”
Workflow encoded in prompts and memory is not a workflow. It is intent. Intent decays. Intent varies by which context loaded in which session. Intent produces the correction patterns that show up in the retro, the ones I spend time fixing instead of building.
A workflow is a thing with a canonical home, a version history, and enough specificity that the agent cannot fill a gap by guessing. It takes longer to write than a memory entry. It survives a context flush. It generates Appendix C.
If a process requires a fresh paragraph of explanation every time someone asks how it works, the explanation is the system, and that system is fragile by design.
Write the contract. Run the reviewers. Promote the document. Then work the 11 items it surfaces. That sequence costs a day. The alternative is paying in small corrections across every session for as long as the system runs.