An unrelated code review turned up a hole that had nothing to do with what I was actually looking at. One query parameter, appended to any page inside an internal admin tool, skipped the staff login check entirely. I confirmed it the ugly way: an anonymous request to a page that should have required a staff session came back with a normal response and live database rows in it.
The mechanism was almost boring once I found it. The staff-auth check ran once, near the top of every admin request, and it had a deliberate early exit for the login page itself, because the login page obviously can’t require you to already be logged in. That exit was keyed off a request parameter rather than the actual page being served. Add the same parameter to any other admin URL and the check waved you through without ever looking at your session.
I fixed it the way you’d expect: stop trusting the parameter, and anchor the exemption to the page the web server actually resolved and served, which a request can’t spoof. Only the real login page kept the exemption. I tested the fix against an unpatched copy side by side. The old bypass URL returned a normal page with data on the unpatched copy and a redirect to login on the patched one. Same request, same account, different code. I shipped it as a hotfix and watched it hold in production.
It held for thirteen hours.
Then staff logins started failing, for everyone, on every environment running the patch. The people who’d never logged out that day didn’t notice, because an existing session still worked fine. Anyone trying to log in fresh got bounced with a generic session-expired error before they ever got a chance to enter a password.
I’d fixed the exemption on the login page, the one that renders the form. I hadn’t noticed that submitting the form is a separate request, a POST to a different internal endpoint, and that endpoint ran through the exact same staff-auth check I’d just tightened. The new guard treated the login submission itself as just another admin request that needed a session it couldn’t have yet, and turned it away before a single credential was checked.
I reproduced it the same way I’d reproduced the original bug: a bare unauthenticated POST straight to the login-processing endpoint, wrong credentials on purpose. Before my fix, that request would have reached the actual authentication check and come back with an invalid-password error. After my fix, it came back with a session-expired redirect instead, meaning the credentials were never even looked at. The guard I’d built to close one door had put a lock on the door people were supposed to walk through.
The fix for the fix was the same principle applied more completely: extend the exemption to the login-processing endpoint too, still anchored to the actual served path rather than to anything a request could set. I verified three things this time instead of one. Wrong credentials against the login endpoint now correctly reached the authentication check and came back invalid, not session-expired. A real login with correct credentials worked end to end. And the original bypass, tested again on the same URL that had leaked data thirteen hours earlier, was still closed. I didn’t want to trade one hole for a different one.
The bypass diagnosis was right. The mechanism was right. Closing it on the resolved path instead of a spoofable parameter was the correct instinct. The mistake was assuming there was one place in the request flow that needed the exemption, when there were two: the page that shows you the door, and the request that actually walks you through it. I tested the first one carefully and never asked whether the second one existed.
I now keep a specific question on hand for any change to an authentication gate: what are every one of the requests a legitimate, logged-out user has to make before they have a session, and does the fix account for all of them, not just the one you were staring at when you wrote it.