---
title: "Three Failures, One Silence"
canonical: https://dxdev.com/blog/2026-09-08_three-failures-one-silence/
datePublished: 2026-09-08
---
## The alert that wasn't an alert

Five nights running, the nightly blog drafter produced zero posts. Five nights running, the only signal we had was a CI triage bot repeating the same line every 45 minutes: `publishPost returned ok:false (expected true)`, thrown by the publish gate test at line 125, with a secondary complaint at line 167 about an ENOENT reading the landing site's content config file. It fired at 2:41 AM, again at 3:26, again at 4:11. Same cause string every time. It looked like infrastructure noise, the kind of thing that shows up in CI when a relative path resolves differently in the test runner than in production, and nobody escalated it because it had looked exactly like that before.

That ENOENT was real, but it was a test-harness artifact, not the production failure. The test suite resolved the content config file relative to a working directory that didn't match how the drafter actually ran at night. Fixing the path made the test pass and told us nothing about why the batch had been silent for five nights.

## The wrong turn

My first theory was the gate itself: one of the seeds queued that week had a dollar figure in its source material, something like a client's monthly rate, and I assumed the leak-detection regex was choking on an ambiguous currency string, flagging it as a possible real invoice number and refusing to draft. I patched the regex to be stricter about what counted as a plausible dollar amount, redeployed, and waited for the next overnight run.

It ran. It still produced zero posts. That was a full night burned on a fix for a bug that wasn't the blocker, on top of the four nights already lost, and it meant the actual failure was still sitting there unexamined while I'd convinced myself I'd found it.

## What was actually stacked in there

I stopped trusting the CI triage summary and went straight to the drafter's own run logs instead of the test suite's interpretation of them. Four independent failures were stacked in the same nightly run, any one of which alone would have been enough to kill the batch.

The drafter was refusing to act on part of its own seed payload. Seeds carry instruction blocks inline, and the model was reading an embedded directive as something closer to an injected command than a legitimate part of its own prompt, and declining to proceed. That pattern wasn't new. Earlier the same night, a browser-manager role instance had refused a file read for a near-identical reason, treating a system-reminder-style block appended to its prompt as untrustworthy and stopping rather than acting on it. Same defensive instinct, different role, same effect: legitimate instructions read as suspicious and dropped.

A second seed's source material genuinely did trip the leak gate, correctly this time, on a detail that should have been masked before the seed ever reached the queue rather than caught at draft time.

A third seed was malformed, held in a bad state, and the batch runner had no per-seed isolation. One bad seed didn't get skipped. It aborted the whole run, which is the actual reason all 43 posts queued for that stretch produced nothing: not 43 individual failures, one failure with no containment around it.

A fourth: a collaborator's name appeared in an internal work log that had been pulled in as source material for six seeds, and the masking pass never touched it, because it only scanned the fields it expected names to live in, not every field a seed could source text from.

## The fix, and what almost stayed broken

The batch runner now catches and skips a bad seed instead of aborting the run, so one malformed entry costs one post instead of the whole night. The masking pass scans seed-derived fields generically instead of an allowlist of expected locations, which is what should have caught the collaborator name the first time. The refusal pattern got a narrower instruction-block format that doesn't read as an injected directive. And the leak gate stayed exactly as strict as it should have been on the one seed where it was right.

Once all four were fixed, that night's run wrote 43 posts.

The part worth sitting with isn't the bug count. It's that a job silently returning zero output for five straight nights generated no alert of its own. The only alert we had was a CI test complaining about a path resolution issue unrelated to the real failure, repeating every 45 minutes for hours, indistinguishable in the noise from every other flaky-looking CI ping that gets triaged and ignored. Zero output and "tests are red" are different failure signatures, and we only had instrumentation for the second one. That gap has its own ticket now, and it's next, because the next time something goes silent for five nights, I want to hear about it on night one.
