At 10:44 PM, session 005 had 33 minutes left. Two entangled work items, a test foundation and a release-tracking board, still had to ship.

The obvious move was to start two build agents and let them race. It would also have been wrong. Both streams touched the same API repository. The board was a prerequisite for the CI gate. The person approving the work was looking from a phone. That is three failed independence tests before the first line of code exists.

I had been treating that judgment as improvisation. By the end of the night, it had become a playbook, a skill, and the first card in a learning ledger.

The work only looked parallel

The initial diagnosis was superficial: two tickets, two deliverables, two agents. That arithmetic is how agent orchestration produces confident rework.

One task needed a test foundation across three repositories, including a smoke test for a legacy ASP application. The other needed a board that showed each repository’s release state and produced a paste-ready repair prompt when one fell behind. The board controlled when the CI gate could become hard. If two builders modified the shared repository while making separate assumptions about that gate, we would merely move the dependency conflict into review.

We took a different path. First, two independent reviewers hardened the plan. They found two real defects: an impossible CI gate in a no-PR workflow, and SHA drift in a deployment script. We also corrected an ownership boundary, retired one misplaced work item, and folded its scope into the right one.

Then the session sent read-only spec agents, one per stream. Each returned an executable spec: exact files and diffs, scope boundary, dependencies, and the decisions that still needed an owner. No spec agent wrote code. The coordinating session did not build either feature.

That distinction mattered because the agent was being a PM, not a faster terminal.

0. Test whether builders are genuinely independent.
1. Fan out read-only spec agents.
2. Commit each spec before synthesizing it.
3. Expose progress in a fixed status pulse.
4. Produce one dependency-ordered backlog.
5. Get one approval, then dispatch builders in order.
6. Run an independent verification pass before review.

The first run exposed a leak. One spec existed only in the coordinator’s context. It was useful until that context ended, then it was effectively gone. A result that cannot survive the session that produced it is not an operating artifact.

So every spec now lands in its work-item folder and is committed before the coordinator synthesizes the backlog. The status signal became a fixed line instead of a prose update:

PM: <n>/<total> specs in (<names>); waiting on <decision>

That pulse gives the team the one fact they need without forcing anyone to reconstruct the session from a chat transcript.

The evidence was in the boundaries

We already had a prioritization rhythm with five beats: intake, prioritize, plan, review, and close. We also had an ICE-C score:

(Impact × Confidence × Conviction) / Effort

Conviction is never inferred. If it is missing, the system asks. Overhead above 20 percent means automate or drop, not simply do it next.

The 10:44 PM session added an orchestration mode. The signal was not that it succeeded once. A lucky one-off should stay a note. The signal was that the sequence had crisp entry conditions, explicit artifacts, a visible failure mode, and a clear reason not to use it in the wrong situation.

We rejected parallel builders because the streams shared a repository, had a dependency chain, and needed a phone-friendly review surface. We rejected letting the coordinator build because coordination state and implementation state blur together. We rejected leaving specs in chat because the first run proved that context-only work disappears with the session. We rejected a new PM console because the existing proposal queue already had the review behavior we needed.

Once I could state those boundaries clearly, continuing to rediscover them in prompts would have been negligent. We wrote the PM playbook and compiled it into a /pm skill with two modes: prioritize and orchestrate.

The skill is not magic. It makes the first decision explicit. Builders are allowed only when the streams have no shared repository, no dependency chain, and a review surface suitable for the person who must approve. If any one condition is false, it goes spec-first.

A skill is not a source of truth

The playbook solved the immediate repetition problem. It did not answer where a learned workflow should live before it becomes a skill.

We had three learning stores with incompatible failure modes. Local agent memories were private to a machine, invisible to teammates and other stations, and already past 280 files. The primary memory file had exceeded its load budget. Skills were versioned and shared inside a repository, but scattered across several repositories with no inventory of what had been codified. The vault was the only store that was shared, versioned, browseable, and reviewable.

We were learning, but learning landed wherever the current session happened to write it. A pattern could exist and still be rediscovered because the next context never loaded that location.

The answer was not to make every observation a skill. We made the vault the learning ledger, and treated memories and skills as compiled projections rather than original sources.

The first pattern card carried the structure we needed:

name: pm-orchestration-loop
status: compiled
proposed_projection: skill
review_tier: notify
source_sessions: [session-005, session-011]
evidence: 2
projection: .claude/skills/pm/SKILL.md

A card is a reviewable claim about a pattern. It includes source sessions, evidence count and links, proposed projection, review tier, and a lifecycle state: candidate, approved, compiled, or retired. Once approved, a compiler can write a memory, create a skill scaffold, add a hook, or schedule a job. Every projection gets a ledger: back-pointer, so a later retro can answer where the behavior came from.

log → mine → card → review → compile → audit

The last step matters. A weekly retro compares compiled behavior to transcripts. A rule that fired and failed gets its card flagged for revision. A rule that never fires ages toward retirement. Self-learning has to include forgetting, otherwise it is just a growing pile of instructions.

Trust needs an inspection point

The ledger also forces a decision about who approves what. We use three review tiers. auto is for reversible, invisible changes that can compile and notify afterward. notify compiles by default and surfaces in a digest. approve is for anything that touches production, money, customer-visible behavior, or agent permissions. Those cards wait in the existing proposal queue.

Directly turning a successful session into runtime behavior would be faster in the narrow sense. It would also make it difficult to see why an agent now behaves differently, and difficult to unwind a bad generalization. The card gives us a pause point between observation and enforcement. The compile step gives us an artifact. The back-pointer gives us provenance. The audit gives us a way to kill rules that have stopped earning their place.

The 33-minute session did not reveal a universal workflow. It revealed a specific one: when tasks look parallel but share code, dependencies, or a constrained review surface, coordination should create durable specs before it dispatches execution.

That is the threshold I am watching for now. When an improvised sequence produces visible state, survives a second use, and has a failure mode we can name, it has stopped being a clever session. It is infrastructure waiting to be built.