Five Alerts, One Ticket
At 8:04 AM the brief fired into Discord five times in under two minutes, all five for the same ticket. That was the flap I sat down to fix. Buried in the noise was a second alarm, quieter but wrong in a more interesting way: a warning that the day’s story points didn’t reconcile with the hours logged against them.
The five-way flap turned out to be the easy part. The brief dispatcher retried on any response slower than its timeout, but it never checked whether the first attempt had actually landed before firing the second. Four of the five messages that morning were retries of a Discord webhook that had succeeded the first time, just slowly. The fix was an idempotency key, the ticket id plus the minute bucket, checked before send. One line of logic, and the flap stopped.
The story-points alarm was the one I should have looked at first and didn’t.
What the alarm was actually comparing
The brief pulled two numbers for the day: total story points closed, and total hours logged by the platform’s dispatch system against those same tickets. If the ratio drifted more than a set band from 1:1, it flagged a mismatch and printed both totals into the 8am brief. Most days it fired. Nobody trusted it, so it got ignored, which is exactly the kind of alert that should not exist.
My assumption going in was the standard one: points are an abstract measure of complexity, hours are wall-clock time, and you should not expect them to converge cleanly. If leadership wanted a real point-to-hour comparison for the brief, the honest fix was to give tickets an actual hours field the dispatch system could write into cleanly, separate from whatever story_points had been holding.
So I wrote it. A migration adding a dispatch_hours column to the ticket table, a backfill script pulling elapsed time from dispatch logs, the whole thing staged and ready to run against three months of history. I ran the backfill on a scratch copy first, the way you’re supposed to, and the numbers it produced were identical to story_points out to two decimal places on every single row.
Not close. Identical.
The column already held hours
That sent me back into the schema instead of forward with the migration. The story_points field was never being written by a human choosing a Fibonacci number off a board. It was being written by the platform dispatch itself, and the value it wrote was minutes worked divided by sixty, rounded to one decimal. Whoever built the original dispatch integration had needed some numeric field to populate automatically and had reused the one already sitting on the ticket schema, because adding a new column felt like more work than it was worth at the time. Nobody had renamed the field, and nobody downstream had ever gone back to check what was actually landing in it.
So the debate the team had been circling for months, whether the platform needed real time-tracking granularity added on top of story points, was already settled. It had been settled since whoever wired the original dispatch integration first reused that field. We just had two names for one number, and had spent an unknown amount of that time treating the mismatch alarm as a data quality problem instead of a naming problem.
I threw out the migration and the backfill script. There was nothing left to build. I renamed nothing either, since a rename would have broken every downstream report reading story_points, and the field was doing its job under the wrong label without hurting anything except my instinct that abstract estimation was supposed to look different from a timesheet.
What shipped
The fix was smaller than the investigation. Every platform dispatch now writes its actual cost against the ticket explicitly, formalizing what story_points had been doing implicitly for however long the integration had existed. The false alarm came out of the brief entirely, since comparing a number against itself was never going to produce a signal. And I wrote up the finding as a standing analysis doc rather than a ticket, because there was no work item to attach it to. The question was closed, not open.
The whole session ran five hours eleven minutes, most of it on a migration that never shipped. The part that actually mattered was running the backfill against scratch data before touching production, because that’s the query that surfaced the identical numbers and stopped me from adding a second column that duplicated a field we already had. If I’d trusted the assumption instead of checking the join, we’d be running two hour-tracking fields today, both correct, both silently disagreeing with each other eventually.