The lease said held. The board said stale. Both were reading the same file, and both were right, for different definitions of “right.”
The immediate trigger was mundane: a teammate’s DNS cutover for a customer domain needed verification (CNAME, apex A, cert, MX, our own domain record), which is routine work I could check in a few minutes. What wasn’t routine was what I found when I went to log it. Another session’s close routine had already run and deleted my session’s active vault manifest, the file that records what a session is doing and lets a human or another agent find it later. I filed a ticket for the deletion itself, restored the manifest at slot 022, and moved on, but only after I noticed. I do not know how long my session would have run with no manifest if I hadn’t happened to go log something right then, and every tool that resolves a session by reading that file would have quietly started failing in ways that look like unrelated bugs, not like a deleted file. That is a worse failure than an error message, because nothing would have told me to look. The near-miss was the interesting part, because it wasn’t a fluke. It was a predictable failure mode of running many concurrent Claude Code sessions against one shared git-backed vault with no lock on manifest deletion.
The setup that makes this possible
The vault is a single git repo. Sessions write session manifests into it, one file per session id, and there’s a pre-close routine that runs when a session ends: check for uncommitted work, check the task board, and, apparently, clean up. That cleanup step didn’t scope its deletes tightly enough. A /close (or a cleanup pass adjacent to it) treated another live session’s manifest as safe to remove, presumably because it looked orphaned by whatever heuristic it used, maybe an age check, maybe a status field, and it wasn’t.
That’s the same failure class you get from desk_authority.py’s lease system, which showed up elsewhere in the same day’s logs. That module manages a single shared “desk lease” so that only one Claude Code session at a time can move browser windows around my screen. Two sessions grabbing the desk to surface a decision page at once would fight over placement, so there’s a lease:
STALE_WINDOW_SEC = 30 * 60
def is_stale(lease: dict) -> bool: age = _age_seconds(lease.get("last_touch", "")) return age is None or age > STALE_WINDOW_SEC
def lease_state(sid: str) -> dict: ... if is_stale(lease): return {"status": "stale", "lease": lease, "age_seconds": age} return {"status": "held", "lease": lease, "age_seconds": age}A lease is held if another session touched it within the last 30 minutes, stale otherwise, and owned if it’s yours. _write_json writes atomically, tmp file then os.replace, specifically so two sessions can’t both believe they hold the lease from a half-written file. That’s the correct instinct: never let concurrent writers race on the same piece of shared, meaningful state. A queue of pending requests sits behind it, so a session that can’t take the lease immediately can ask and wait rather than force it.
The manifest deletion had no equivalent guard. Nothing staled it, nothing queued a request to remove it, nothing atomically checked “is this actually orphaned” before unlinking it. It just went.
Alternatives considered
Lock the manifest directory like the desk. Reuse desk_authority.py’s lease pattern: before any close routine deletes a manifest, it has to claim a lease on that specific sid’s file, which fails if the file was touched inside the stale window. This is the fix I’d converge on, since the mechanism already exists and is proven; it’s a matter of pointing it at manifests instead of only at desk placement.
Never delete, only archive. Have close routines move manifests to a closed/ subdirectory instead of deleting them. This is simpler and safer by construction: even a wrongly-triggered cleanup just misfiles something instead of destroying it, and it’s a two-line change. The tradeoff is it doesn’t stop the underlying bug, a session’s manifest can still vanish from where every tool expects to find it, active, mid-task. It’s a good second layer, not a substitute for the first.
Git as the safety net. The vault is versioned, so in principle nothing is unrecoverable, restoring slot 022 was a git checkout on that path once I noticed. But this only works if you notice. The failure mode here wasn’t data loss, it was silent state corruption: my session kept running with no manifest, which meant every tool that resolves a session by reading its manifest would have started failing in ways that look like unrelated bugs. Git being able to restore the file after the fact doesn’t help the stretch of time, however long it actually was, where nothing knows the file is missing. I’m keeping git as the last-resort backstop, but it’s not the fix.
What ships
A ticket is filed for the deletion itself. I’m extending the desk lease’s stale-window pattern to manifest writes: before any close routine deletes or overwrites another session’s manifest, it takes a per-file lease keyed on last_touch, same 30-minute window, same atomic-write discipline that already protects the desk. Given how many sessions are running against this vault at once on a given day, that path needed the same care as the one that stops two agents from fighting over a browser window.