---
title: "The Auto-Heal Worker That Almost Broke Itself"
canonical: https://dxdev.com/blog/2026-08-22_automation-that-ate-its-own-state-file/
datePublished: 2026-06-26
---
At 2:54 p.m., the live production log showed the certificate auto-heal worker trying to process `cf-heal-state.json` as a command.

That file was not a command. It was the worker's cooldown ledger, the record that stops it from re-triggering the same certificate repair over and over. I had written it as JSON because it was structured state. I had put it in the queue directory because that was where the worker was already writing files. The queue scanner treated every `*.json` file in that directory as work.

So the worker's memory became work. The system did not break a customer site, but it had built the exact loop that could make a healer chew on its own state every time the scanner ran.

## The job began with a real HTTPS failure

We had a migrated customer site stuck without HTTPS because its certificate had not completed. I manually unstuck that certificate, then audited all 1,282 hostnames rather than assuming it was an isolated failure.

My first instinct was too broad. I re-triggered moved hostnames before putting a proper DNS gate in front of that action. The audit made the problem visible, but it also showed why an unbounded repair script was the wrong follow-up. A certificate re-trigger is not an operation to spray across every hostname that looks adjacent to the incident.

The system we shipped had three separate layers instead. The domain dialog showed the true certificate state. Staff got a one-click re-trigger control for a known bad state. A background worker then handled certificates that were genuinely stuck, so recovery did not depend on a customer reporting a broken site.

The worker had precision guards: a pointed-to-us gate, a cooldown, and a per-pass cap. The first guard kept it from acting on hostnames that did not point to us. The second kept it from retrying the same repair continuously. The third bounded the amount of work a single pass could do.

That was the real architectural decision. A manual button alone loses because it keeps recovery reactive. Re-triggering every moved hostname loses because it ignores the conditions that make the action safe. The worker was the right path only because its permissions were narrow enough to match the failure.

## One glob was the command interface

The deployment used a file-backed queue. The worker scanned the queue directory with `*.json`. Any JSON file there was treated as a command payload.

That made the filename extension part of the protocol. It was not a cosmetic choice.

`cf-heal-state.json` lived in the same directory. To the code that wrote the cooldown ledger, it was a state file. To the queue scanner, it was a command. Both interpretations were locally reasonable. Together, they were a namespace collision.

The diagnostic path mattered here. At first, the fault looked like it could be in the certificate flow: an API response, a bad hostname classification, or the cooldown logic itself. The worker had just gone live after a production certificate incident, so each of those was plausible. The production log showed the actual shape of the problem. It was not failing to remember a cooldown. It was reading the cooldown ledger through the wrong entry point and processing it as a bogus command.

That distinction changed the fix. I was not debugging JSON parsing. I was debugging an interface boundary that existed only because two unrelated files shared a directory and an extension.

## State cannot look like work

The repair was small. The cooldown ledger stopped being a `.json` file in the queue namespace. The durable convention was simpler: queue commands are JSON, cooldown ledgers use `.txt`.

A named exclusion for `cf-heal-state.json` would have patched this instance, but it would have left the bad boundary in place. The queue would still mean both “incoming commands” and “whatever support files happen to be stored here.” The next state file would have needed another exception.

A separate state directory would also solve it, and it is the cleaner choice when the worker has more state than a small cooldown ledger. In this case, changing the ledger's type was enough to make the contract explicit without changing the command scanner that was already working.

The important part was not the suffix. It was restoring one meaning per intake rule. `*.json` was allowed to mean “process this” again because the worker no longer wrote its own JSON state into that scope.

## The worker needed a boundary before it needed more intelligence

The auto-heal worker exists because a stuck certificate should not wait for someone to notice that HTTPS is missing. But remediation code is unusually good at creating secondary failures. It runs after something has already gone wrong, it has permission to make changes, and it often needs persistent memory to avoid doing the same thing twice.

That makes its storage choices part of the control plane. A cooldown ledger is not an implementation detail when it decides whether the worker acts. A queue directory is not just a folder when a glob turns filenames into commands.

I had given the worker a cooldown so it would not repeat repairs. Then I put the cooldown file where the worker looked for repairs. The live log caught the contradiction before it turned into a larger incident.

The final fix changed one file type. The real fix was giving state and work different names, different meanings, and no opportunity to impersonate each other.
