---
title: "We Paid For Our Most Expensive Model To Fix Workers That Were Never Broken"
canonical: https://dxdev.com/blog/2026-09-19_paid-to-fix-a-problem-that-did-not-exist/
datePublished: 2026-09-19
---
Three transcripts ended on a clean `stop_hook_summary` followed by a prompt nobody had answered. That was enough for us to declare three lane managers dead and respawn all three on our most expensive model. Twenty-some minutes later, two of the originals ran another clean turn.

## What "dead" looked like

Each manager had dispatched child work and ended its turn. After that, silence. The tail of every transcript was the same: hook summary, then the last prompt with no reply. No child results came back either. It looked like a crash, so we treated it as one and started fresh sessions on Opus.

That was the wrong call. We paid for three new top-tier sessions to replace three sessions that were fine.

## Checking the bodies

The first hypothesis was capacity eviction. The bridge's `makeRoom` only evicts idle sessions, skips any with an attached watcher, and heals itself on the next turn through `--resume`. I grepped the 15 daily delivery ledgers for the two strings it would leave behind:

```bash
grep -riE "killed or evicted|bridge at capacity" .state/deliveries/*.jsonl
```

Zero hits across all 15 files. Eviction was never in play.

Next, the transcripts themselves. All three were intact and resumable at 924K, 1.3M and 1.2M. Every turn each manager had been given ended with `terminal_reason: completed` and `is_error: false`, with no `api_error` and no quota notice. Then the timestamps: two of them ran another clean 5-hook turn at 15:35 to 15:36Z, more than 20 minutes after we had buried them.

The signature that fooled us, a clean stop-hook summary and then an unanswered prompt, is ordinary CLI trailer bookkeeping. A healthy idle session has it too. It has no diagnostic power on its own.

## Why idle and dead look identical from outside

A bridge session is a transcript id, not a process. Every turn is a fresh spawn:

```
["-p", prompt, "--model", opts.model || "sonnet"]   // plus --resume <fullId> on later turns
```

The process table agreed: six `claude.exe -p ...` processes, all parented to the bridge host, one of them carrying `--resume`. The process that answered turn N is gone before turn N+1 is dispatched.

Only two code paths inject a message into an existing session, and both need an HTTP request from a human or an agent. No timer, queue or completion handler wakes anything. A manager that dispatches work and ends its turn goes inert until something outside speaks to it. From the outside, "waiting correctly" and "stuck" produce the same bytes.

## The instruction that made it worse

The `/manager` skill told managers to arm a background waiter before ending a turn. Under this process model that cannot work, because the waiter is a child of a process that exits when the turn ends. Each of the three transcripts carried six "Background shell command didn't finish before the previous session ended" notices, the predictable result of following the instruction. That instruction was the actionable bug. The sessions were healthy.

## Nothing durable to read back

Even a manager that is alive has little to check. `.runs/agents/<sid>.jsonl` records that a subagent finished and never what it returned. Across 3,277 all-time `done` rows, the longest `detail` field is 172 characters. Opt-in status is thinner still: of 34 sessions on the day, 9 wrote any `.status.jsonl` row and 4 emitted a terminal `done` or `failed`, about 12%. So the completion event gets recorded, the content lives only in a transcript, and the manager has almost nothing durable to poll.

## The premium that mostly did not apply

The respawn had a second flaw. `opts.model || "sonnet"` means any resume turn whose sender omits a `model` field falls back to Sonnet, and the helper we use to answer a session sends no model field. A manager spawned on Opus silently drops to Sonnet from its second turn, and the transcript flags nothing. So the money went to top-tier first turns on sessions that then ran on the cheaper model, replacing sessions that never needed replacing.

## What we changed

Three fixes, all pending as hand-ups:

- A durable per-lane board that the child writes at its own completion, with no opt-in.
- Dropping the background-waiter instruction from `/manager`, since it cannot survive the process model.
- Stamping `model` on every resume turn.

The check itself is cheap. Open the transcript, read the last turn's `terminal_reason`, and confirm the file still resumes. A session that completed its last turn and resumes cleanly is idle, not dead, and we no longer respawn it.
