The alert wasn’t an alert. It was a number: 136 rejected sends in the accounting logs over one week, and nobody had looked at what was actually in that pile.
We run mail through our MTA behind the platform, and rejections get logged the same way bounces do, envId keyed, one CSV per day. I’d already been burned by trusting a single day’s accounting file after an earlier fix: the report that joined order rows to MTA accounting by envId only read one day’s file, so anything still retrying at midnight rolled into the next day’s CSV and rendered as “inQueue/noData,” which staff read as “never sent” when it had actually gone out and bounced. That fix taught me not to trust a rejection count from a single file or a single query. So when 136 showed up, I didn’t ask “how many bounced,” I asked a session to pull the raw accounting rows and join them back to real orders before anyone touched a keyboard about it.
The join
The session read the week’s accounting CSVs, matched envId back to order and customer records, and split the 136 into two piles: 99 were internal batch mail, staff notifications and system chatter, nobody outside the company ever saw them. The other 37 were customer-facing, spread across roughly 25 distinct addresses. Within that 37, it broke out by message type and found expiry notices were the single biggest category: renewal and lapse warnings that customers never got, meaning a chunk of those 25 people had a site heading toward expiration with no warning delivered. That split came straight out of the join, not out of a hunch. Some customers had multiple failed sends stacked on the same address, which is why 37 rejected messages maps to about 25 people, not 37.
That’s the part that made this different from a routine cleanup ticket. A batch-mail rejection is noise. An expiry notice that silently failed is a customer finding out their site is down when it’s already down.
Where I got it wrong first
My first instinct was to have the session draft a single blast apology, one template, mail merge across all 25 addresses, get it out same day. I actually had it draft that version. It came back technically fine: correct tone, correct facts, a generic “we noticed a delivery issue affecting your account” line. Then I read the underlying rows again next to the draft and saw the problem: eight of those addresses belonged to accounts with prior support history on the same thread, in some cases with names attached in the message body of earlier exchanges. A form email to those eight would have read as a company that doesn’t know its own customers, on the exact subject where that matters most. I killed the blast draft. That was a real wrong turn, not a hypothetical one. I had the session ready to queue it before I caught it in the diff.
So the session redid the split: two customer-facing apology templates, one generic for the 17 with no prior thread, and it named the 8 accounts specifically that needed a human-written reply instead of a form letter. It drafted both templates. It did not send either one, did not queue anything in the mail system, and did not touch any mail-server configuration, no retry settings, no queue priority, nothing. It left the ticket open with the drafts attached and a note listing exactly what it had and hadn’t done.
The second pass
Before we acted on any of it, I ran a second, independent session against the same raw CSV, no shared context, no pointer to the first session’s output, just the source files. Its job was to recompute the counts and the address list from scratch: 136 total, 99 internal, 37 customer-facing, roughly 25 addresses, expiry notices as the largest customer-facing category, same 8 accounts flagged for a human reply. It matched, row for row.
That’s not caution for its own sake. A rejection count that’s off by even a few changes who gets a form letter and who gets a personal one, and getting that wrong in the wrong direction is its own kind of damage. Running the same join twice, blind, cost one extra session and confirmed the numbers were real before either drafted email went near a send queue.
Nothing about the analysis needed me in the loop. The session could read the accounting files, do the join, count the categories, and draft the language faster and more carefully than I would have working it by hand at 6 AM. But the drafts sat in the ticket, not the outbox. Autonomy got the counting right twice. It didn’t get anywhere near the send button once.