---
title: "Alert on Delivery Outcome, Not Queue Depth"
canonical: https://dxdev.com/blog/2026-09-26_mail-queue-alert-outcome-vs-infrastructure/
datePublished: 2026-08-16
---
The alert landed at 12:53 on a Saturday, and it said ERROR six times.

```
5 Min Monitor: ERROR - THERE ARE ISSUES!
MTA Queue: ERROR
Item had High value: 1762
Item had High value: 1446
VMTA Servers: ERROR
(1762) newsletters
Domains: ERROR
(1446) freemail-a.com
(311) freemail-b.com - Error: "552 1 Requested mail action aborted, mailbox not found" ... [2026-08-15 12:52:38]
Relayed: 89% good (290 bounced out of last 2652 today)
Mailers: 100% good (0 had errors out of 66 in the past 60 minutes)
Emails: 100% good (0 were invalid out of 104 in the past 60 minutes)
```

I read it the next morning and spent 38 minutes on it. Nothing was wrong.

## What the monitor actually measures

We run a commercial MTA, version 4.0r19, on an unmanaged Linux box, split into four virtual MTAs, one per sending IP. The `newsletters` vmta carries bulk runs. The 5-minute monitor polls each vmta's queue and each destination domain's queue, and it trips when a count crosses a fixed threshold.

A large newsletter send makes exactly that count. The 1,762 queued on `newsletters` was almost entirely the two domains named below it: 1,446 for the first freemail provider plus 311 for the second is 1,757. That is a batch sitting in the queue while the MTA works through it at the pace each receiving server allows. The queue is supposed to look like that mid-send.

The alert bundles two different kinds of check. "MTA Queue," "VMTA Servers" and "Domains" are infrastructure state: how many messages are waiting. "Mailers" and "Emails" are closer to outcome: did requests fail, were addresses invalid. Both outcome lines said 100% good. Only the depth lines were red.

## The wrong turn

I did not start with that reading. The second provider's line has an SMTP error string in it, and we had a real delivery incident this summer. For roughly ten days, the DNS record behind one sending IP's reverse-DNS pair was missing from our DNS host, and that provider rejected the stream with `550 5.7.25`. That one was invisible until someone noticed mail not arriving.

So I went hunting for a repeat. I searched our notes for the monitor's name and got seven stale reconciliation-queue history files, nothing about the monitor. I read the forward-confirmed reverse DNS section of our internal infrastructure inventory, which covers all four vmtas and which zone each A record lives in. I was checking whether the `newsletters` PTR still resolved back to its IP.

That was the wrong question, and it took a good part of the 38 minutes to see why. The error is a `552 ... mailbox not found`. That is a recipient-level hard bounce: the receiving server was reached, accepted the connection, and said this one address does not exist. A reverse-DNS failure looks different. It is a `550 5.7.25` at connection time, and it hits every message on that IP rather than one address at a time. A `552` is what dead addresses on a newsletter list produce.

The inventory section was useful reading and it cleared the DNS question, but I only got there because I had matched an error string to the wrong incident. The alert's own shape should have sent me to the bounce line first.

## What settled it

Two facts closed it:

1. The outcome lines were clean. The `290 bounced out of last 2652` line reads as a worrying 89% good, but the errors on it were recipient-level `552`s, not the connection-level `5.7.x` rejections that would point at the IP.
2. The MTA drained the queue on its own. No restart, no intervention. The depth came back down after the batch finished.

Conclusion: queue-depth thresholds tripping during normal delivery. No action needed.

## What I'd change in the monitor

The threshold is a static count on a quantity that legitimately swings depending on whether someone pressed Send. A count tells you how much work exists. It says nothing about whether the work is failing. The alerts I'd keep page on outcome:

- **Bounce rate over a trailing window per vmta**, with the receiving domain's rejection code attached. A jump in `5.7.x` policy rejections pages someone. A scattering of `552 mailbox not found` does not.
- **Age of the oldest queued message**, not the number of messages. A queue of 1,762 that is all under a few minutes old is a healthy send. A queue of 40 that has been stuck for two hours is a real problem, and the current monitor would say OK to it.
- **Queue depth as context on the alert, never as the trigger.** It is a useful number to read once you are already looking.

The reverse-DNS incident argues for this too. That failure would never have moved a depth counter much. It would have shown up as a rejection pattern across one vmta, which is the thing worth watching.

## The cost of the current setup

The direct cost was 38 minutes on Sunday for a Saturday alert. The larger cost is that a monitor which goes red on every big send teaches us to read its ERRORs as noise. When the next real rDNS-style failure arrives, it will look like every other red alert we ignored on the way there. Alert fatigue is a delivery risk of its own.

The fix is small, and we haven't made it yet: keep the depth numbers on the report page where they can be read, and move the paging to bounce rate and oldest-message age.
