At 10:47 PM I sat down with a GPT-authored strategy document open on one screen and the live vault open on the other, and started marking the document wrong, one recommendation at a time. Six hours and eleven minutes later I’d rewritten most of it, written down a canonical ticket workflow that had never existed on paper before that night, and found two real bugs in how our agent sessions put finished work in front of me to review. None of those three things were the plan when I opened the laptop.

What the strategy doc actually got wrong

Earlier that night, at 1:42 AM, a different agent had already reviewed the same document and returned a conditional verdict: approve the measurement direction, but do not proceed dashboard-first or checklist-first. That’s a sequencing call, not a rejection of the whole plan. The document’s instinct, build something visible to show progress, was right. Its order of operations was backwards.

The GPT plan wanted a dashboard first, something you could point at and show someone. What actually needed to happen first was checkout-date tracking and a reliable snapshot underneath it. We’d already shipped the checklist and checkout-date hotfixes days earlier, and the nightly metrics snapshot had only just gone live that night. Build a dashboard on top of that and every number on it means something. Build it first and you’ve shipped a page that visualizes nothing.

So I went through the document against the vault and my own history of past ticket decisions, correcting it recommendation by recommendation, and wrote down for the first time what my actual stance is on agent-system investment: fund the plumbing before the display. That’s sitting in the vault now, for the next session that has this same argument with itself.

Then I tried to just look at the page

With the sequencing settled, I built the front end: a Conversion page and release markers, on top of the snapshot data that was now real. Standard move to close the loop: point the agent browser at the new page, capture it, and have it waiting for me.

It came back blank. A page shell, no chart, no rows, exactly what you’d see if the metrics endpoint had never been called. I re-ran it. Blank again. I assumed it was a data problem and checked the nightly snapshot job directly against its output in the database. The data was there.

The actual cause was in the browser session, not the page. Earlier that same night I’d fixed the agent browser to sign in to platform staff automatically, so sessions wouldn’t need to be walked through login by hand every time. That auto sign-in was writing its session cookie against a fixed QA tenant regardless of which environment the target page actually lived on. The Conversion page was rendering under a test account that had never run the checklist flow, so the query it fired came back empty, correctly. The login shortcut that was supposed to save time was quietly authenticating every capture against the wrong tenant.

I fixed the host binding and reran the capture. A second bug showed up immediately: the screenshot fired before the page’s async metrics fetch had resolved, so now I had a page correctly authenticated against real data, captured mid-spinner. That one was a straight race. The browser tool took its shot on page-load, not on network-idle, and a chart that fetches after mount loses that race close to every time.

Neither bug had anything to do with conversion metrics logic. Both live in the layer that takes what a session built and puts it in front of a human to look at, and neither had ever been caught by a unit test, because nothing about them is a unit. They only show up when a real page, with a real chart, gets rendered through the actual login path a real session uses. I filed both as follow-up tickets instead of patching them inline, because the auto sign-in host binding affects every other page any agent might screenshot for review, not this one alone.

What stuck

The GPT document wasn’t wrong to write. It got a plausible plan down fast and gave a second agent and me something concrete to disagree with, which is faster than starting from nothing. But the value wasn’t in the document itself. It was in what happened after: correcting it against real history produced a written policy that hadn’t existed before that night, and the unrelated task of just getting a finished page onto a screen surfaced two bugs that had nothing to do with conversion metrics at all. I went looking for a dashboard and came back with a rule about session hosts and a rule about screenshot timing. Both are more durable than the dashboard was.