---
title: "My Fix for a System Freeze Caused Its Own 16-Minute Freeze"
canonical: https://dxdev.com/blog/2026-08-22_the-fix-that-deadlocked-itself/
datePublished: 2026-04-10
---
Two hard freezes in roughly two hours turned a 7.5 GB always-on box into an unavailable agent runner. The immediate temptation was to call it a memory leak. It was not. The machine was being asked to commit more memory than it could safely carry while cron jobs launched overlapping Claude subprocesses.

I built a spawn lock to stop that. Then the lock turned a heavy cron job into its own 16-minute freeze.

That second failure was less dramatic than a hard system lockup, but it was more embarrassing. I had added a mechanism whose entire job was to make concurrent work safe. Instead, I put the same exclusive lock on both sides of a parent and child process boundary, then waited for it to resolve itself.

It could not.

## The failure I was actually trying to stop

The first incident was a capacity problem with a scheduling problem layered on top. The box was small. Claude subprocesses were useful, but they were not cheap, and cron could start a new one while another one was still active. Two hard freezes made the risk visible. The useful diagnostic was not a steadily growing process that would point to a conventional leak. It was chronic `Committed_AS` overcommit.

That distinction changed the remediation. Finding and killing a leaking process would not solve a system that could be healthy one minute and then accept one concurrent memory commitment too many the next. The fix had to constrain both the amount one subprocess could reserve and the number of subprocesses allowed to begin together.

I added more swap, per-process `RLIMIT_AS`, and a `MemoryMax` control at the cgroup level. Those controls cover different failure modes. Extra swap gives the machine more room to absorb a spike. `RLIMIT_AS` prevents one process from reserving an unlimited address space. `MemoryMax` gives the workload a harder container-level ceiling.

None of them is a scheduling primitive.

The system still needed a universal spawn lock across all 12 known places that could start a Claude subprocess. That was the part that sounded simple. Open a lock file. Take an exclusive POSIX lock. Start the subprocess. Release the lock when the critical section ends.

The mistake was deciding where that critical section ended.

## A lock at the wrong depth

The first version put `fcntl.flock()` in the cron wrapper. The wrapper acquired the lock before it invoked its child. The child process then entered the normal Claude spawn path, which also acquired the same lock.

The shape of the bug was effectively this:

```python
# cron wrapper
with open(lock_path, "w") as lock_file:
    fcntl.flock(lock_file.fileno(), fcntl.LOCK_EX)
    subprocess.run(["python", "worker.py"])

# worker.py
with open(lock_path, "w") as lock_file:
    fcntl.flock(lock_file.fileno(), fcntl.LOCK_EX)
    start_claude()
```

That looks like defensive duplication if you read the two files separately. It is a deadlock when you read them as one process tree.

`flock()` is not a reentrant mutex shared by a parent and its child. The parent owns the exclusive lock. It waits for the child to finish. The child waits for the parent to release the exclusive lock. Neither condition can become true first.

A heavy cron job sat there for about 16 minutes before I noticed and killed it. The original machine freeze had made me add a control that could block launches. The first version of that control had no escape route once its own process tree disagreed about who owned the lock.

## Why the obvious alternatives were not enough

It would have been easy to stop at the memory controls and call the incident closed. That would have reduced the chance of another hard freeze, but it would have left 12 separate spawn sites free to create the same burst pattern. More swap could buy time. It could not serialize an unsafe launch sequence.

It would also have been easy to leave the lock in the wrapper and remove it from the worker. That would have worked for the cron path and quietly weakened every other call site. A universal rule implemented in only one caller is not universal. The callers change. New scripts appear. Someone starts the subprocess from a place that never touches the wrapper.

A global lock was still the right control. The wrapper was the wrong place to own it.

## The corrected boundary

The repaired design moved lock acquisition to the leaf-level spawn function, the one function that actually starts the Claude subprocess. The parent can orchestrate work, schedule it, retry it, or collect its output. It does not take the spawn lock. The leaf takes the lock immediately before the expensive process begins and releases it when that launch path is finished.

That made the invariant narrow enough to audit:

> Every Claude subprocess launch passes through one leaf-level function, and only that function acquires the cross-process spawn lock.

I also added an in-process reentrant guard. The POSIX lock handles contention between processes. The guard handles accidental nested calls within the same process. They solve different problems, so neither replaces the other.

The final layer was a hard 15-minute cron timeout. That is not a cure for deadlock. It is an admission that any control which can wait must have a bounded wait somewhere above it. If the same class of bug returned through a future refactor, the cron job could fail visibly instead of holding the scheduler indefinitely.

I also kept the circuit breaker that pages before the next hard freeze. The point was not to build one perfect protection. It was to make each failure mode fail earlier and more legibly.

The incident changed how I read a lock. The question is not only, “What resource does this protect?” It is also, “Which process owns it, which descendants need it, and what work is the owner waiting for?”

A parent lock around a child that needs the same lock is not extra safety. It is a process-tree cycle with a filename attached.

The spawn lock is still there. It now protects the one transition it was meant to protect, and it can no longer block that transition by protecting it twice.
