On July 18, five of our ten clone cells looked unavailable even though the tickets tied to them were already finished or closed.
That was enough to stall new work before anyone had written a line of code. We run ten parallel clones of the same repository so developers and agent sessions can take separate tickets without grinding on the same checkout. The useful abstraction is not just a clone. It is a clone cell: a checkout, the ticket it is serving, and the branch that may still contain work.
At that scale, an abandoned claim is not an edge case. It is a capacity leak.
Occupied is not the same as active
The first symptom was a pool that looked full. The wrong response would have been to wipe the status board, mark every clone free, and start dispatching again. That would recover apparent capacity quickly, but it would treat every old claim as dead code. One of the ten cells was carrying an in progress prototype branch, and another feature branch was local only and not backed by the remote. A blanket clear would have converted a scheduling problem into a recovery problem.
The opposite mistake would have been to trust the existing occupancy marks. That keeps the pool safe, but it quietly turns finished work into permanent reservations. Neither option distinguishes a stale claim from work that merely stopped at a good breakpoint.
So I audited all ten cells against their tickets. The result split cleanly into three groups. Five cells were attached to work that was finished or closed. Three were on genuinely open tickets. One prototype branch was deliberately parked on its clone. After freeing the five stale or done occupancies, six cells were available for new work.
That count matters. It is the evidence that the cleanup changed usable capacity, not just labels in a dashboard. Three active cells, one parked prototype, and six open cells account for the full pool of ten.
The branch changes the answer
A ticket is the first diagnostic, not the whole diagnostic. A closed ticket often tells you the application work is complete. It does not tell you whether a checkout contains something that only exists there.
The local only feature branch was the useful failure mode. Its ticket state made the cell look like a candidate for release. Its branch state meant that clearing it without a decision could discard the only copy of unfinished work. The follow up was explicit: push it to the remote if the work needed safeguarding, otherwise leave the audit closed. That is a much better question than, “Is this clone old?”
The sequence I want now is simple:
- Start with the pool, not the session that happens to be noisy.
- Match every clone cell to its current ticket state.
- Separate finished or closed work from open work.
- Check whether an apparently stale cell has a branch that still needs a home.
- Free only the cells whose ticket and branch state agree.
That is intentionally narrower than automating a universal cleanup command. An automatic age threshold would be attractive, especially when an agent session dies without updating its claim. It would also be wrong for the parked prototype. Time is a signal. It is not authority.
Capacity is a coordination problem
The early mental model was that more clones meant more parallelism. It is true, but incomplete. More clones also mean more state that can become inaccurate. A clone does not stop being occupied because an agent went quiet. It stops being occupied when we prove that its ticket, branch, and intended next action no longer need that cell.
This is why I treat the pool audit as operating work, not housekeeping. It took 4 hours and 59 minutes to audit all ten cells and free five of them. That is not glamorous work, but it is what made six new lines of work possible without destroying a prototype or losing a local branch.
Ten clones do not give us ten times the throughput by themselves. They give us ten places where a stale claim can hide. The capacity comes back only when we clear those claims with evidence.