The p90 queue delay on our password reset email was 697 seconds during a campaign burst and 3 seconds outside one. Users had been telling us that resets and signup confirmations felt slow. Nobody had a number until a hub session of mine, an AI agent I use as a daily working surface, went looking for one.
One channel, three kinds of mail
Everything the product sends goes out through a single virtual MTA on a single IP. That covers password resets, signup confirmations, and the bulk mailers we run for win-backs and announcements. A queue like that is FIFO per channel. It has no idea that one message is a person staring at a login screen and another is the last item in a promotional batch.
When a campaign starts, the bulk messages land ahead of whatever urgent mail arrives next. The urgent mail waits its turn. At the p90, its turn came 697 seconds later. Outside a burst the queue is nearly empty and the same messages leave in about 3 seconds. That is why the problem was so easy to miss. The average looked fine, the quiet-hours behavior looked fine, and only the moment a campaign kicked off made anything look wrong.
The fix is to give the two kinds of mail separate channels, so bulk volume can never sit in front of a reset. I filed that as a ticket, wrote the decision up in the vault, and moved on to the harder question of why it had taken this long to see.
The plan that was already on the books
An older ticket proposed a routing change as the answer to “email is slow.” Once the 697 versus 3 split was on the table, that plan was dead. Splitting urgent from bulk is a different shape of change from rerouting, so I marked the routing ticket abandoned in favour of the new priority-channels one.
Leaving it alone has a cost. A dead ticket that still looks alive keeps getting rediscovered. Any session that searches for email routing finds the old plan, and nothing in it says it was superseded. Each rediscovery is a session, agent or human, evaluating a plan we had already thrown out. So I wrote up the decision and the reason for it, with a pointer from the old ticket, to make the abandonment the first thing anyone finds. Closing a ticket is not the same as recording why it is closed.
The next burst was already loaded
While I was in the queue data, I looked at the campaign that was next in line: a win-back mailer to 4,085 addresses. It was configured to go out on the transactional IP, the same one carrying password resets.
That was a few hours after we had shipped a List-Unsubscribe headers hotfix for that mailer and fixed two template bugs. A test send had landed in spam. If we had sent it as configured, we would have manufactured exactly the burst that produced the 697-second p90, and we would have done it on the channel that also holds our sender reputation for the mail people actually need to receive. A bad batch to a cold list can drag a reset email into the spam folder too.
So the From address has to change before that send. We also ran an outside review of the plan, and it corrected our bounce thresholds sharply upward. A list that old will bounce far more than one built from recent signups, and a threshold tuned for a warm list would have tripped a stop condition almost immediately. The review also flagged reassigned role inboxes as the real blind spot. An address like a shared team inbox can be handed to someone else years after it was collected, and it looks valid to every check we had. Nothing was sent. The resume order is on the ticket.
What made this visible
Nothing here needed new tooling. It needed someone to ask what shares a channel with what, and to look at delay at the p90 during a burst instead of the mean over a day. Shared infrastructure keeps working right up until the moment two workloads want it at once, and until then every dashboard reads green.
If you run email through one channel, pull the queue delay for your transactional mail and split it by whether a bulk send was active. If the two numbers match, you are fine. Ours were 697 and 3.