The ticket closed clean: diff applied, tests green, no complaints from review. Except when I went back to check the worker session log for that ticket, there wasn’t one. The front desk had marked it done, but nothing in the session registry showed a worker ever spawned for it. The only session that existed was the front desk’s own.
That’s the bug. The secretary substituted itself for the worker.
The architecture, briefly
The front desk is the orchestrator in this system. It doesn’t write code. Its job is to take a ticket, look up which role should handle it, attach to that role’s running session if one exists, or spawn a fresh one if it doesn’t, and then drive that session turn by turn: feed it the ticket context, forward its questions, collect its output, close the loop. Roles are locked down with their own MCP server declarations and permission scopes, which is the whole point of separating orchestrator from worker: the front desk gets broad visibility, the worker gets a narrow, auditable slice.
The dispatch code looked roughly like this:
def dispatch_to_role(role, ticket): session = registry.find_session(role) if session is None: session = spawn_role(role, ticket) return drive(session, ticket)spawn_role is supposed to be the fallback when no session exists yet. For that ticket, the role in question, one we’d only ever driven interactively before, had no registered spawn path. spawn_role hit that missing path, caught the failure internally, and returned None instead of raising. drive(None, ticket) then fell into a branch that had been written for local testing months earlier: if there’s no session object, just run the ticket’s tool calls directly in the current context. That branch was never supposed to survive into production. It did, quietly, because nothing ever exercised it until a role without a spawn path showed up.
Chasing it down the wrong way first
My first fix was wrong, and it cost me an afternoon. I assumed the problem was a timing race, that find_session was looking before the worker had finished registering itself, so I added a retry loop: three attempts with backoff before falling through. It seemed to work. Tickets for that role stopped erroring immediately.
But the fallback branch was still there, just harder to reach. On the next ticket that hit a role with no spawn path, the retries burned nine seconds, timed out, and then quietly dropped into the same self-execution branch as before, just later and buried under three retry log lines instead of one clean failure. I spent that afternoon reading through longer logs convinced I had a flaky registry lookup, when the actual defect was that the self-execution path existed at all. The retry logic didn’t close the gap, it just made the gap take longer to fall into and harder to spot in the transcript.
What actually fixed it
The real fix was to delete the fallback and make the absence of a spawn path a hard failure:
def dispatch_to_role(role, ticket): session = registry.find_session(role) if session is None: session = spawn_role(role, ticket) if session is None: raise NoWorkerSession(role, ticket) return drive(session, ticket)No silent substitution, no partial success. If a role can’t be spawned, the ticket fails loudly and shows up as a registry gap to fix, not as a mystery diff with no worker log behind it.
Why this is the diagnostic that matters
The tell isn’t an error message. It’s a ticket that closes cleanly with no corresponding worker session. That’s the signature of an orchestrator quietly doing the job itself instead of delegating, and it will pass every functional test you have, because the output is fine. The only way to catch it is to check that the session registry, not just the ticket outcome, matches what you expect. If your secretary can always produce a correct result even when the delegation path is broken, you don’t have a multi-agent system. You have one agent wearing a title, and the failure only shows up when you go looking for the worker that was never there.