---
title: "Snapshot Drift: When Your Team's Clone Falls Behind"
canonical: https://dxdev.com/blog/2026-09-26_snapshot-staleness-team-dev/
datePublished: 2026-08-13
---
`Invalid column name 'referer'` was the yellow error box on my teammate's admin schedule page. The same URL rendered fine on my machine. Same commit, same branch, same code path. The page does an INSERT into `dbo.staffactions` on staff login-as, and on his machine that INSERT died.

## The column that only live had

The error names a column, so the code and the database disagreed about the schema. I had the AI session query the SQL Server directly. Live had `referer` on `dbo.staffactions`. The dev database and all four dated snapshot databases did not.

A change had shipped on Aug 6 that logs the referrer on the staff login-as path. It added `referer varchar(255) NULL` to live at deploy time. The dev database and the snapshots were copied before that date, so none of them ever got the column.

My teammate's clone had `DB_OVERRIDE` in `.env-local` pointing at one of the snapshots, a copy from late July. All of my clones have that variable blank, which means live. That was why it worked for me. The override was the only difference between us.

Nothing announced this. There was no migration runner for snapshot databases, no check that a clone's schema matched what the code expected, and no warning on startup. The first signal was a crash on a page that only staff open.

## The fix that patched the wrong thing

The session's first recommendation was to backfill the column into the old copies:

```sql
ALTER TABLE dbo.staffactions ADD referer varchar(255) NULL
```

Run that against the dev database and the snapshots, leave the override alone, and the page works. It was a nullable add on non-production databases and looked cheap and safe. The session also argued that a new person shouldn't be on the live database anyway, and I went along with it.

I asked for a paste-ready prompt to hand to my teammate's agent. The prompt it produced was careful. It read the override value and refused to continue if it was blank. It checked `sys.columns` for `referer` and expected 0. It ran the ALTER only if the check returned 0, then re-ran the check and expected 1. It ended with an explicit instruction not to change the override, because staying on the snapshot was intentional.

Then I read it again and asked whether it made him stop pointing at a legacy database. It did the opposite. I had asked for a fix, taken the first one on offer, and had it written up as an instruction, without stopping to ask what setup I wanted him to end up in. I wanted him to match every other clone I run, which is override blank and pointed at live. The prompt I'd been handed patched a stale copy so he could keep using it.

That cost me a full round trip and a second prompt written from scratch. The new one told his agent to read the current override value and report it, then set the line to exactly `DB_OVERRIDE=` (blank, not deleted), recycle the local IIS site, and reload the schedule URL to confirm it rendered. It also said that from then on he was on the shared production database, where reads were fine and any direct write was a stop-and-ask.

The session also flagged the one thing I couldn't answer from the logs. If someone had parked him on a snapshot on purpose, so a new person couldn't hurt prod while learning, then blanking the override removes a guardrail. That is a question for whoever set up his machine. I made the call to move him to live, but I had to make it deliberately. The first prompt would have made it for me by default.

## The ticket for the next clone

Moving one teammate off a stale snapshot fixes one clone. It doesn't fix the next person who ends up on a snapshot-backed clone, and every one of those will hit the same crash the moment they trigger a staff login-as. So I filed a follow-up to backfill `referer` into the dev database and the snapshots, the same ALTER the first prompt contained. It was a good fix for the wrong problem. As a follow-up ticket it earns its place.

## Snapshots sit still while the code moves

A snapshot is a copy of production at one moment, and the code keeps moving. The column that killed the page was added in a small, correct hotfix. Nobody did anything wrong. The code and live moved together, and the snapshots sat still. The gap was about three weeks wide by the time someone stepped in it, and the only detector was a person opening the right page.

Two cheap checks would have caught this before any human did:

- A startup assertion in dev that compares the columns the code writes against `sys.columns` for the database it's actually connected to, and prints the override value when they disagree.
- A rule that any schema change to live also lands in the dev database and the current snapshot the same day, or the change doesn't ship.

I haven't built either yet. The first is about twenty lines. The second means someone owns the snapshots, which is a harder problem than the assertion. For now the override line is the first thing I ask anyone about when a page works for me and crashes for them.
