524s on the main site, measured over five complete days: 6 on the 24th, 9 on the 26th, 1 the day after. The Dev Notes on the ticket had called it one account’s data hitting a slow path. That reading died the moment the flood that supposedly caused it was gone and the timeouts kept showing up on separate days anyway.

Cloudflare gives an origin 100 seconds before it gives up and returns a 524. Over that five day window the main site logged 40 failed requests total: 24 were 522s on / and /teams/default.asp, 16 were 524s. Same family, different half of the same fuse. A 522 means the connection to the origin died before the 100 second ceiling. A 524 means the origin was still working when the ceiling hit. Both point at the same thing: something on the other end is slow, and Cloudflare doesn’t care why.

A companion ticket had already done the work of measuring why. /teams/default.asp -> home.asp on an org site was pulling miniScheduleWide at 11,682ms, versus 330ms on a plain league site, a 35x gap for what’s nominally the same widget. Roster and standings were 3.9x and 4.5x slower on the org path. The query count on that page dropped from 255 to 123 once the fix landed, and a helper called InactiveExludeSQLGet that had been firing 34 times for 1,479ms got memoized down to one call. That fix was real, it was measured, and it was still on develop, not master. So every number on the timeout ticket was a baseline that was about to move under it. Re-running the Cloudflare counts against that ticket before the fix shipped would have meant tuning against a floor that was already scheduled to drop.

When I went through the pages on the timeout list, they split into two categories that no amount of query tuning touches the same way.

Four of the worst offenders had Server.ScriptTimeout set higher than Cloudflare’s edge ceiling:

  • ScheduleWizard.asp: 180s
  • Registration_FileUpload.asp: 600s
  • MonitorReport.asp: 2400s
  • MailerOrderSend.asp: 2700s

ScheduleWizard.asp alone accounted for 7 of the 24 measured timeouts in the window. Across that whole set, the server was configured to let the script run for anywhere from 3 minutes to 45 minutes past the point where Cloudflare had already decided to hang up and return an error. It doesn’t matter if the underlying query gets 4x faster. Shave MailerOrderSend.asp from 2700 seconds down to 900 and it is still 9x past the 100 second ceiling it will die against on the edge. That’s not a query optimization problem, it’s an async/chunking problem: the work has to move off the request/response cycle entirely, either queued and polled or broken into pages, because no amount of speeding up a synchronous script gets it under a ceiling it was never designed to respect.

The other bucket is /admin/home.asp, pageedit.asp, and teaminfo.asp, which came in through a separate ticket reporting IIS killing scripts at the origin rather than Cloudflare timing out at the edge. Those pages don’t set an inflated ScriptTimeout. They pull SportsHQ-s.asp, Prefs-s.asp, and Rst-s.asp through a shared Common-s.asp include, the same include chain the org-site fix touched. Nobody had measured those three pages directly yet, but they sit on the exact code path that was already getting rewritten. That bucket is a real candidate for query optimization to just work, because it was never architecturally doomed the way the four long-timeout pages are.

We merged two tickets on this into one because they were the same failure measured from two ends of the same connection: IIS killing the script at the origin is one ticket’s view, Cloudflare giving up waiting for that same script is the other’s. Worth keeping straight when you’re staring at an error rate dashboard: a 522 and a 524 in the same week aren’t two bugs, they’re one slow origin observed from both sides of a 100 second wall. And a page whose timeout is configured above that wall isn’t a performance bug at all. It’s a page that was never going to succeed, dressed up as one that just needs to try harder.