---
title: "The Provider Saying Yes Wasn't the Same as the Job Being Finished"
canonical: https://dxdev.com/blog/2026-06-17_the-provider-said-yes-before-we-were-actually-done/
datePublished: 2026-06-17
---
# The Provider Saying Yes Wasn't the Same as the Job Being Finished

The idea was simple: when an automated job finishes, send a real text message to a person, not a quiet log entry nobody's watching or a green light in a browser tab that's already closed. A short message that reaches someone even when they've stopped paying attention to the run.

Building that during a production cutover, I believed the hard part was already handled. Move the sending duty to the new worker carefully, verify it, keep the old path listening as backup just in case. That plan solved the problem I was watching for. It did not occur to me that the messaging queue itself, the thing responsible for carrying that safer version of the send, had its own separate risk sitting underneath, until I sat with the failure mode long enough to ask a more specific question: what happens if the process handling this crashes at the exact moment between the message going out and our own system recording that it went out?

That gap is easy to miss because both halves of it sound like "sent." The outside messaging provider accepting a request is one fact. Our own system finishing its bookkeeping and marking the job complete is a separate fact, a moment later. If a crash happens in between those two moments and the job retries, from the system's point of view nothing has been recorded yet, so it tries again. The provider accepts the second attempt too. The same message goes out twice, to a real person, for a mistake that only ever happened in the space between two events that felt like they should have been one.

A single marker written only after everything finished wouldn't close that gap, because the crash happens before that marker ever gets written. The fix needed two separate signals: a short-lived claim set before the outside request goes out at all, so a near-simultaneous retry doesn't double up on something already in flight, and a second, longer-lived marker set only after the provider actually confirms it received the message. Two facts, recorded at two different times, because they are genuinely two different events separated by a real window where a crash can land.

## If something in your process talks to an outside service

Look specifically at the space between "the outside service said yes" and "my own system wrote down that this is done." If those two things happen in one step, with one marker set only at the end, a crash in between can make a retry repeat the exact thing you were trying to make safe. A message resent, a charge run twice, a notification duplicated. The fix isn't more retries. It's a marker on each side of that gap, so a retry can tell the difference between "never happened" and "already happened, just not filed yet."
