The console popped up first. The crawl script was shelling out to git with no window-hiding guard, and every so often a terminal would flash on screen mid-crawl. That was the bug that got fixed at 9:07 PM. The crawler itself came two hours later.
The idea was simple: before anything ships to production, walk every admin page with a staff-authenticated session and see what breaks. Not a unit test, not a fixture, an actual headless browser hitting actual staging URLs in sequence, the same way a real login would. We already had a headless-guard test in the suite for exactly the console-leak class of bug, so once the git-shell issue was caught it was fast to verify: run the existing guard against the fixed script, confirm no spawned window, commit, push.
The crawler itself took about two and a half hours to get right. First pass authenticated once, then round-tripped through the sitemap and hit every admin route with a plain GET, logging status codes and any 500s. That caught nothing interesting, everything came back 200, because a 200 on a page with a broken form action tells you nothing about whether the form actually works. The whole reason for building this instead of leaning harder on the existing test suite was that the 6,000-test suite I run pre-close was already a known pain point that day (a separate ticket existed specifically because pre-close was running the full suite instead of a targeted subset). Status-code crawling was the version of “automated” that gives you a green check for nothing. I ran it anyway, got a clean report, and almost called it done there.
It wasn’t done. A clean crawl on GET requests doesn’t tell you if a batch operation is wired to the wrong table. So the second pass walked deeper: instead of just requesting each page, it looked for the actual regression class we cared about, which was catchable by inspecting the response body for known failure signatures (missing table data, broken redirects, stale references) rather than just the HTTP status. That’s what an AI code-review pass had flagged the same evening on a related admin batch-archive change: new tables were being copied and remapped, but the per-sport delete-original step wasn’t wired to match, a correctness bug that would return 200 all day while quietly leaving orphaned data behind. Same shape of problem: nothing throws, nothing 500s, it just does the wrong thing under a green light.
The rebuilt crawl found it. One real regression, filed as a single ticket, in staging, before it touched production. Everything else the crawl flagged turned out to be pre-existing broken admin pages that had been broken for a while and nobody had noticed because nobody was walking that part of the site regularly. Those got filed separately so they didn’t drown out the one that actually mattered.
The whole run, first crawl to second crawl to filed ticket, took two hours and thirty-eight minutes, from 8:59 PM to 11:37 PM. After that it got scheduled: a weekly Tuesday check, same authenticated crawl, same signature-based inspection instead of status-code-only, running against staging before anything moves to prod.
What made the difference wasn’t the browser automation, that part’s plumbing. It was the shift from “did the page load” to “does the page do the right thing,” which meant reading response bodies for the specific failure signatures a real regression leaves behind instead of trusting a 200. The first version of the script would have run every week and reported all-clear on a bug that was actively going to ship. The second version is the one that’s still running.