---
title: "The Bottleneck in the Middle"
canonical: https://dxdev.com/blog/2026-09-26_blog-pipeline-end-to-end/
datePublished: 2026-09-25
---
At 12:19 AM the cron-fail triage agent read a failed dxdev run and answered in 9 seconds, 451 tokens: `spend_gate deferred the dxdev model call (band orange, discretionary), not a code bug.` The job had not crashed. It had been told, correctly, to wait.

That one line summarizes what our blog pipeline was like before we fixed it, and what we changed.

## Five stages that failed in combination

The pipeline had five stages that each worked in isolation and failed in combination: weekly mining of work logs for post candidates, description generation, lesson extraction, spend gating, and a dash linter for the finished drafts. We fixed all five in one pass instead of one a week, because fixing them separately kept moving the bottleneck to the next stage down.

After the pass, 27 posts went live, and the nightly job drained the rest of the backlog without anyone touching it. We had spent months treating the backlog as a writing problem. It was a plumbing problem. Drafts existed. Nothing between the draft and the publish step was reliable enough to run unattended.

## A deferral filed as a crash

Mining and publishing are the ends. Everyone builds those first, because they are the visible parts: a thing goes in, a post comes out. The middle is where a model call has to be allowed to happen, where its output has to be judged, and where a failure has to be told apart from a decision.

The 12:19 alert is that last part. Our spend gate sorts model calls into bands. At band orange, anything marked discretionary is deferred. A blog call is discretionary, so at orange it waits. The job wrapper saw a call that did not complete and filed it under `remedy/cron-fail`, the same bucket a real crash lands in. A deliberate deferral was being reported as a failure.

The triage agent got it right, but only because it read the gate's reason string. Anything that just counted failed runs would have called that night broken. The cost of the mislabel is real, even if small: every deferral generates a triage call, and every triage call is a model call. We were spending money to be told that we had decided not to spend money.

## A lessons stage allowed to say nothing

At 3:40 AM the `learning_nightly` job ran twice, 13 seconds apart. Both times it returned a single line, `VERDICT: SKIP`, at 78 and 71 tokens. Nothing was wrong. Both times the stage looked at the day's material and decided there was no lesson worth keeping.

That is a stage designed properly. Earlier versions of lesson extraction always produced something, because we had written the prompt to require output. A stage forced to emit a lesson every night emits filler on the nights with nothing to say, and filler is what makes a backlog of posts read as generated. Giving the stage a legal way to say nothing is what let the rest of the pipeline trust it.

The gate needs the same thing: a legal way to say "not now" that downstream stages read as a state, not an error. The fix on the alerting side is small. When the gate defers, record the reason and the band, and have the wrapper exit with a distinct status so it never enters the failure channel at all.

## Why nobody watched the middle

The ends of a pipeline get attention because a human touches them. I read the mined candidates. I read the published post. Everything in between is machinery that nobody watches when it works, so nobody notices it is held together by assumptions: that the budget is always green, that every stage always has something to say, that a non-zero exit always means a bug.

Each of those assumptions was false at least once, and each one had the same signature. A stage did the correct thing, and the layer around it reported the correct thing as a fault. Once the fault reports were trustworthy, the backlog drained on its own. Nothing about the mining or publishing changed for that to happen.

If your automated job has a human at the start and a human at the end, look at the middle first. Read the reason strings in your failure channel for one week. If some of them are decisions, your pipeline is already working better than its dashboard says, and the stall you are trying to fix is in how the pipeline reports itself.
