The admin error dashboard showed 352 next to one page. The ticket that got filed off that number described it as 352 times in the past week. That reading made it sound like something had just started breaking badly. It hadn’t.
What the number actually was
The query behind that count had no date filter in it at all. It was a straight lifetime total, grouped by error type, with no window on it whatsoever. What made it misleading is where it sat: directly beside a “most recent occurrence” column showing a timestamp from that same day. A large number next to a recent timestamp reads as “this many, just now,” even though the two columns were answering completely different questions.
Querying the raw error table directly instead of trusting the dashboard’s framing told the real story: that count’s actual spread went back more than two years, to a specific month in 2024. Spread across that whole span, it was a slow trickle. Not nothing, but nowhere near the emergency the ticket had been filed as.
The fix, and why it had to be structural
Explaining the real numbers once wasn’t enough, because the dashboard would keep producing the same misleading pair for the next person who looked at it. The fix added a second column, a genuinely time-bounded count of occurrences in just the last seven days, computed with the same query using a conditional sum rather than a separate round trip to the database. The old lifetime column got relabeled “All time” so it could no longer be mistaken for something current. One real page on the live dashboard now reads 142 all-time next to 6 in the last seven days, and anyone glancing at it can tell instantly which number matters for right now.
A second bug, found while looking
The same ticket also turned up something concrete in the underlying error data: a handful of upload pages were timing out, and the pattern in the log was one account retrying the identical file upload roughly ten times over two days, never once succeeding. The setting that limits how long a script may run was missing from three of these pages, while thirteen comparable pages already had it set. That setting covers more than server-side processing time. It covers the entire time the server spends receiving the file from the browser, which for a normal phone photo on a slow connection can run well past a minute on its own. Without the setting raised, that upload dies mid-transfer every time, and because transfer time is just a function of file size and connection speed, the exact same file fails the exact same way on every retry, no matter how many times someone tries again.
Confirmed the cause with a direct comparison on a test server: the identical file, over an identical slow connection, failed with a server error when the setting was unset and succeeded once it was raised to match the other pages. Fixed all three.
The actual lesson
A number with no time window attached will always get read in the context of whatever else is on the screen next to it. Here that context was a recent timestamp, so a backlog more than two years deep read as an active flood. The fix wasn’t just correcting one misfiled ticket. It was making sure the dashboard itself could no longer produce that misreading for the next person who glances at it.