---
title: "When Executors Lose Auth Context"
canonical: https://dxdev.com/blog/2026-09-26_executor-privilege-context-breaks-auth/
datePublished: 2026-09-23
---
Twenty-six rewritten posts were staged, the stuck promote branch was merged, and nothing published. The run had gone a full pass with no error in the output that anyone read.

The blog lane runs headless. A scheduled task launches the executors, and one of them is `codex exec`, which does the scoring and rewrite passes. On this run its auth was broken for most of the pass. We later traced that to the executors running as SYSTEM.

## The wrong turn: I believed the queue

The session before this one ran 15 hours. It fixed four bugs that were quietly blocking the pipeline, backfilled the missing lesson blocks across the whole review queue, and wrote nineteen drafts for the gap week on both the technical and non-technical tracks. Its conclusion was that nothing was publishable for content reasons rather than plumbing. It also noted that part of the scoring record looked untrustworthy.

I took the first half of that at face value. The morning's blog run started from it. I merged the promote branch and staged the 26 rewrites on the assumption that the plumbing was sound and the content was the problem, so the rewrites were the fix. That cost a full pass of rewriting posts against a verdict that came from a scorer I now had reason to doubt. The run took one minute, and every rewrite in it was aimed at a diagnosis I hadn't checked.

I can't prove the scoring gap and the auth failure are the same event. Both showed up in the same window, and the pipeline reported neither as a failure.

## What SYSTEM does to a CLI login

An interactive `codex` login writes its credentials under the profile of the user who ran it. A scheduled executor running as SYSTEM gets a different profile directory, so it looks for that file and finds nothing. On Windows the exact paths differ by service context, but the mechanism is that the login exists for one identity and the job runs as another.

The interactive shell works, because it is the identity that logged in. The scheduled job fails, because it isn't. That is why nothing looked wrong when I ran things by hand.

The failure was quiet for a second reason. The executor's output lands as a one-line summary in the session log, and a job that cannot authenticate produces a summary that looks like any other short one. Elsewhere in the same day's log, headless runs that hit a timeout at 900 seconds record `(timeout after 900s)` with no token count and no cost. Those are at least visibly wrong. An auth failure that returns quickly and empty has no such marker.

## The fix, and the check I should have had

We fixed it by making the executors run with credentials that belong to the context they run in. The pipeline logic stayed untouched. After that the executors could authenticate.

The check I would add is a preflight that runs before any lane does real work. It makes one cheap authenticated call as the same identity the lane will use, and it stops the run loudly if the call fails. I'd put it ahead of the scoring step, because a lane that stages 26 posts and publishes zero should read as a failed run, and right now it reads as a quiet one.

The other change is smaller. I no longer treat a "nothing is publishable" verdict as a content finding until I have confirmed the scorer that produced it could authenticate when it ran. A scoring record from a job that may have been running blind is not evidence about the posts.

## Where the run stands

The 26 staged rewrites are still staged. Whether they were the right rewrites is an open question, because the verdict they were built on came from a pipeline that could not have been reliably checking anything. Until the preflight exists, a green run here only tells me the job finished, not that it did its work.
