I found the failure during a production messaging-worker cutover, before I ever trusted SMS as the way an autonomous job talks back to a human. A BullMQ worker can hand messages.create to Twilio, crash before the queue marks the job complete, and then send the same verdict twice when it retries.
The original problem was a production messaging-worker cutover. It became the design for something more useful: a callback that tells me when an autonomous job is actually finished, what it decided, and what I need to do next. Not a green toast in a browser tab. Not a record buried in a queue dashboard. A short text that reaches me when I am no longer watching the run.
That distinction sounds obvious until a job succeeds in the narrowest possible sense. A worker can finish its computation, write a success record, and disappear into the background. The human who started it still has no answer. Worse, a failed notification can look like a successful job if the UI only reports what the worker tried to do.
Once the worker-first cutover order was decided, I believed the hard part was over. Move the drain path, verify the new worker, leave the old API listening as a safety net. That plan solved the loopback problem. It did not occur to me that the queue itself, the thing carrying the safe version of the send, had its own duplicate-delivery risk sitting underneath, until I sat with the failure mode long enough to ask what happens if the worker dies between Twilio saying yes and BullMQ saying done. I had already told myself the cutover was handled before I found that question.
We traced the wrong path first
During the cutover, we traced the old send path. An older server-side app made a loopback request directly to localhost:3000/api/notify/send. The IIS proxy rule existed, but it only handled inbound web traffic. The loopback call bypassed it completely.
That changed the risk calculation. Stopping the old API before the replacement bound port 3000 would not create a graceful blip. It would create a synchronous, unretried loss window. The caller would attempt the notification, find nothing listening, and move on.
The first obvious answer was a hard stop, start the new service, and rely on the proxy. We rejected it because the proxy was not on the call path that mattered. The safer cutover was worker-first: leave the old API listening so work could still be enqueued, move the drain path, then verify the new worker before touching the port.
That diagnostic mattered for autonomous jobs too. A notification path has to be evaluated from the point where the job emits it, not from the browser or public route that happens to look healthy. If the worker is making the call, I test the worker’s path.
Provider acceptance is not a completed callback
Moving the send into BullMQ fixed the hard stop problem, but it exposed a subtler one. BullMQ is at-least-once. If a worker stalls after Twilio accepts the request but before BullMQ records completion, stall recovery can run the job again. The job did not fail from the queue’s perspective. It became eligible to send the same text twice.
For an autonomous callback, duplicates are more than annoying. A second text saying “deploy verified” after a rollback, or a repeated request for approval, makes the human doubt the result. I needed the callback to be idempotent at the provider boundary, not merely idempotent when it entered the queue.
I considered three shortcuts. A dashboard badge lost because it only works while I am looking at the dashboard. A Redis flag checked when the job starts lost because a crash can happen later, immediately around the provider call. Writing the durable marker only after messages.create returned lost because it left the duplicate window open.
The pattern that survived was a short-lived claim before the provider call, followed by a longer-lived sent marker after Twilio accepted it:
const claimed = await redis.set( `agent-callback:${runId}`, "pending", "NX", "EX", 600);
if (!claimed) return;
await client.messages.create({ to: operatorPhone, from: callbackNumber, body: verdict});
await redis.set( `agent-callback:${runId}`, "sent", "XX", "EX", 86400);The ordering is the contract. SETNX claims the right to send. Twilio receives messages.create. Only then does the claim become a 24-hour sent marker. The 10-minute pending TTL must be shorter than the BullMQ stall-recovery delay. If the process dies after the claim but before Twilio accepts the message, the pending key expires before the retry runs, and the retry can send. If the process dies after Twilio accepts the message, the sent marker blocks the duplicate.
A hard provider failure follows a third path. We delete or tombstone the claim so a legitimate retry is not suppressed. Treating every failure as sent would hide the exact callback the human needed.
The callback needs its own receipt
The text itself is not the end of the diagnostic trail. Twilio’s status webhook posts back to /api/notify/webhook/status, where we record the provider’s delivery state against the callback. That gives the job two separate facts: Twilio accepted the request, and the callback later reached a terminal delivery state.
I do not make the autonomous job wait for a final handset receipt before it completes. Mobile networks are not a transaction coordinator. The job completes after it has made the callback durable and handed it to the provider. The delivery update is an observable follow-on state. If delivery fails, the system can escalate or surface that failure without pretending the underlying operation failed.
The test contract is deliberately ugly. We inject a crash after SETNX but before the provider call, then confirm the retry sends. We inject another crash after provider acceptance but before BullMQ completion, then confirm the retry does not resend. We also verify that the status webhook can update the callback record independently of the worker.
The two numbers that make the pattern hold are a 10-minute pending TTL and a 24-hour sent marker, in that order, on either side of one Twilio call. Get the ordering or the TTL wrong and the pattern collapses back into exactly the bug it was built to close: a worker that dies at the wrong instant, and a human getting told twice that a rollback happened, or not at all.