---
title: "The Self-Updating Deploy Script That Has to Fail Once to Fix Itself"
canonical: https://dxdev.com/blog/self-updating-deploy-script-fails-once-to-fix-itself/
datePublished: 2026-06-01
---
I pushed a one-line fix to a deploy script, the deploy ran, and it failed at the exact line I'd just fixed. Not a different line, not a new error, the same line, with the same message I had spent the last ten minutes patching. For a second I was sure I'd botched the fix. I hadn't. The fix was already on disk. The script just wasn't running it yet, because it updates its own source in the middle of its own run and keeps executing the copy it loaded a commit ago.

If you have a self-hosted runner that pulls fresh code as part of the deploy, this is a trap waiting for you, and the symptom is a bad one: a correct fix that produces an identical failure. Here's the whole thing.

## The original failure

The job was "Deploy API" on GitHub Actions, and it went red in 37 seconds. The main API service itself started fine and passed its health check, so this wasn't the app breaking. The job died later, in the deploy script's service-management step, at the start helper, line 72:

```
Cannot find any service with service name 'worker-queue-api'.
```

The deploy script (`deploy.ps1`) walks a list of services, stops them, swaps the build, then starts them back up. That list is the main API service plus the two SMS-queue services. The two SMS-queue services aren't actually provisioned on this box yet, so stopping and starting them is asking Windows for services that don't exist.

The root cause was an asymmetry between the stop and start helpers. The stop helper guarded against missing services:

```powershell
$svc = Get-Service -Name $name -ErrorAction SilentlyContinue
if (-not $svc) { Write-Warning "$name not registered (skipping stop)"; return }
```

The start helper did not. It went straight at the service and threw hard when it wasn't there. So the deploy would happily skip stopping the two phantom SMS-queue services, then blow up trying to start them. The fix was obvious: mirror the same guard into the start helper so a not-yet-provisioned service logs a warning and is skipped instead of killing the run:

```powershell
$svc = Get-Service -Name $name -ErrorAction SilentlyContinue
if (-not $svc) { Write-Warning "$name not registered (skipping start)"; return }
```

One file, one change. I committed the fix, pushed to `main`, and watched the run.

## The twist

It failed again. Same job, same start helper, same `Cannot find any service` message at the same line. My change was sitting right there in `main` on GitHub, and the runner was producing output as if it had never seen it.

The instinct here is to assume you fixed the wrong thing. I almost re-opened the diff to hunt for a second copy of the bad code. Don't. Look at the order of operations on a self-hosted runner that resets its own working tree.

Here's what actually happens. The runner is itself driven by `deploy.ps1`. When the job starts, PowerShell parses the entire script into memory from whatever is currently on disk, which is the *previous* commit, because the tree hasn't been updated yet this run. Then, partway through that already-parsed script, there's a step that does:

```powershell
git reset --hard origin/main
```

That command updates the file on disk to the new commit. But the script that's running is the one PowerShell already parsed at the top, the old one. The start helper executing at line 72 is the old version, without the guard, even though the file on disk now has the guard. The script reset its own source out from under itself and kept running the version it loaded before the reset.

So the fix was in. It just wasn't the one running.

## Confirming it instead of assuming it

The tempting move here is "that must be it, it'll pass next time" and a blind re-run. I didn't want to re-run on a theory. I SSH'd into the box (the self-hosted runner, which runs on the prod machine) and read the on-disk `deploy.ps1` directly. The guard was there. The `git reset --hard origin/main` from the failed run had already pulled the fix onto disk. So the on-disk script was correct, and the only reason the last run failed was that it had been parsed before that pull.

With the file confirmed, I re-ran the deploy. Green. The log showed exactly what I expected: `not registered (skipping start)` as a warning for each phantom SMS-queue service, and `Health check PASS`. The health probe only ever hit the main API service's health check endpoint, so skipping the two missing SMS-queue services didn't weaken what the deploy actually verifies. I checked the run output with `gh run view --log`.

Two runs, one fix. The first run failed *because* it was the run that installed the fix.

## The real fix

Any deploy script that does `git reset --hard origin/main` (or any in-place self-update) partway through its own execution is running the previous version of itself on the run that performs the update. Interpreted languages parse the whole script up front, so the change you pushed lands on disk during the run but doesn't take effect until the run after.

The cleanest fix is to kill the surprise entirely. Reorder the script so the self-update happens before anything load-bearing, then re-exec the freshly-pulled script in a child process:

```powershell
# At the very top of deploy.ps1
if (-not $env:DEPLOY_REEXECED) {
    git reset --hard origin/main
    $env:DEPLOY_REEXECED = "1"
    & $PSCommandPath @args
    exit $LASTEXITCODE
}
```

`& $PSCommandPath` launches a fresh PowerShell process from the on-disk file, which is now the new version. The env flag prevents the loop. After this, the version executing and the version on disk are always the same commit, and a script can never fail just because it was the run that fixed itself.

## Related

- [One Private Dependency, Five Different Failures](five-layer-deploy-break-one-private-dep): another multi-layer deploy fragility on the same self-hosted pipeline
- [Our grandfathered legacy GitHub plan silently blocked every CI run, and the cheaper fix was a one-way door](github-bronze-legacy-plan-silently-blocks-all-ci): a deploy environment assumption that silently fails for the same structural reason
- [My red CI was a lie: a deleted workflow haunted every push while nothing real ran](ghost-workflow-zero-second-failure-masked-no-ci-running): another pipeline state where the failure indicator lags one run behind reality
- [My pinned app kept vanishing after every reboot, and it wasn't Windows being flaky](pinned-app-vanishes-after-reboot-self-updater-staging-folder): self-updater behavior causing the same one-run-behind surprise on a different platform
- [One monorepo, two build lanes: keeping classic-ASP pushes at zero CI minutes](monorepo-two-build-lanes-zero-ci-minutes-legacy-spa): structuring the pipeline so deploy scripts and build artifacts stay in separate lanes
