---
title: "An Error Spike Was Actually Two Separate Problems"
canonical: https://dxdev.com/blog/2026-08-16_two-incidents-in-one-error-spike/
datePublished: 2026-08-16
---
A little after 8pm, the client-side error log lit up. New error types, several different ones, all appearing for the first time that night, clustered in the same forty-minute stretch. That pattern usually means one thing: something just got deployed and it broke.

Nothing had just been deployed.

## Two things, not one

Pulling the actual error detail split the spike into two unrelated causes. The bulk of the volume, maybe eight different error types, traced to a single visitor on an eleven-year-old browser. Its JavaScript engine couldn't even parse one of the site's script files, so every function that file was supposed to define came back undefined, and every click on the page threw a different reference error. One stuck device, retrying the same broken action for forty minutes. Not a regression. Not something to chase.

The second cause was smaller in volume but real: a crash in the schedule page's tournament filter, hit by several distinct visitors on current browsers, going back at least four days before that night. It had simply been happening quietly the whole time, and one of its occurrences happened to land inside the same window as the noisy browser.

Only one of those two things was a bug worth fixing. Filing that apart from the noise was the first thing that mattered.

## Filed at the wrong priority

The filter crash got a ticket and went into the normal queue, on track to ship with the regular weekly release. That felt like the right call in the moment: it wasn't a new outage, it had been going on for days already, and a queued fix seemed proportionate.

Asked directly whether all of tonight's findings should ship as emergency fixes, the honest answer was no for most of them. But not for that one. A defect already confirmed hitting several real users repeatedly, for four days running, doesn't get to wait behind the routine queue just because it wasn't discovered as an emergency. It was rebuilt as an immediate fix instead, merged, deployed, and confirmed live on the real page with a cache-busting check before calling it done.

## Following the spike to its actual shape

With the real bug shipped, the spike itself still needed an explanation: why had a genuinely rare crash and one dead browser combined to look like a flood at 8:50pm specifically? The site's own error monitor counts rows in a rolling thirty-minute window and pages someone once that count crosses a threshold. Querying the same window it uses showed the count climbing from 24 errors at 8:30pm to a peak of 68 to 69 around 8:50 to 8:55pm, then falling back to single digits by 9:25pm, exactly the shape one stuck browser retrying for forty minutes would produce layered on top of ordinary background noise.

## What the rest of the sweep turned up

Understanding the shape of tonight's spike didn't mean the error log was actually clean. Working through the rest of it, and checking each candidate against the tracker to confirm nothing already covered it, turned up five more genuine, previously unfiled bugs: a role-management dialog that crashed for a class of staff accounts, a sponsors page erroring on an empty background selector, a schedule wizard misrouting to a handler that had been deleted years earlier, a chat widget wiring itself up before the feature flag that gates it was actually ready, and a leftover call into an embed integration that no longer needed it. All six, counting the filter crash, shipped as immediate fixes before the night was over.

## The actual takeaway

A crash cluster is not automatically one incident just because it arrived at the same time. Splitting signal from noise first meant fixing exactly one thing that mattered instead of chasing a dead browser. And once that one thing was confirmed as a real, recurring, multi-user bug, its priority should have come from that fact on its own, not from which pass of triage happened to turn it up. The queue a bug lands in should track how much it's already hurting people, not how it was found.
