At 4:56 AM the triage job named a cause for a red build: nanoid <3.3.18, a high-severity DoS, GHSA-2v37-7h3g-55p8, fixed in 3.3.18. The dependency path in the alert was sanitize-html > postcss > nanoid. We don’t depend on nanoid directly. It sits two levels down under a sanitizer, which is why a scanner flags it and a grep of our own package.json doesn’t.
At 8:26 PM the same advisory fired again, this time worded as an infinite loop, same path, same fixed version. That is 15 and a half hours of a known high-severity finding sitting in the queue. I have no story about a recut advisory. The record shows the same GHSA id and the same 3.3.18 floor both times, and a patch that only shipped at 8:52 PM. What I can tell you is why the gap existed: it was our alerting, not the dependency.
Three bad alerts before the good one
The night before, the pipeline had a different problem. At 1:26 AM a deploy alert came in with a cause line reading “health check failed after deploy AND automatic rollback also failed; app may be down.” The autofix step answered CANNOT-FIX, because the failing step was only the deploy workflow’s failure-reporting step. At 1:56 AM another alert arrived. This time triage said “cause unclear” and autofix said the failing-step log was unavailable, so there was nothing to diagnose.
Those were the two shapes of failure: a real cause hidden behind the reporting step, and no log at all. So when the nanoid line finally appeared at 4:56 AM, my first instinct was to distrust it. Three alerts in a row had named the wrong thing or nothing. I let it sit while I dealt with other work.
Repairing the messenger while the advisory waited
That was the wrong turn, and it had a price. I spent the early afternoon repairing the alert path instead of reading the one alert that was accurate. A red build had named the wrong cause again, so I changed triage to read the step that actually failed rather than the last step in the workflow.
The repair tooling had its own costs, and I only found them because I was in there:
- The auto-repair tool kept its working copy after every run, parking 426MB each time. Clearing it reclaimed 1.4GB in total.
- The test suite for the alerting spawned real repair jobs against live state. Tests were triggering the thing they were meant to test.
- A set of collector deploy scripts existed only in a scratch folder, uncommitted.
Every one of those was worth fixing. None of them was the advisory. The nanoid finding was correct the first time it fired, and I treated a correct alert as noise because its neighbors had earned that.
Patching a dependency you never declared
The 8:26 PM alert was the one I acted on. The patch went out as a normal hotfix, to develop and to main, and I confirmed the deploy came back up before calling it done. Because the vulnerable copy is transitive, the check that matters is the resolved tree, not the manifest. The version you declare and the version that resolves under sanitize-html are separate facts, and only the second one is exposed.
When a triage bot has been wrong three times in a night, the fourth alert still gets read against the advisory database, not against the bot’s track record. GHSA-2v37-7h3g-55p8 takes thirty seconds to look up. I spent an afternoon on the plumbing instead, and the high-severity finding stayed open until evening.