---
title: "A Working Credential Is Not the Same as a Working Agent"
canonical: https://dxdev.com/blog/2026-08-21_a-working-credential-is-not-the-same-as-a-working-agent/
datePublished: 2026-08-21
---
I had a rule for who a message should come from when something needed a person's attention right away. It should go out under that person's own assistant, the one built to act on whatever the message was about, preferring it over a shared company identity that could only relay the words. The rule I wrote to enforce that was simple: if that assistant's own credential worked, send under its name.

The credential is just an API key. It works whether or not anything is listening on the other end. I could post a message as that assistant with the process that was supposed to read it and act on it sitting completely dead, and the send would still succeed. The system would report the right identity delivered the message. Nobody would know the one thing that identity was actually for, someone or something being awake to respond, had not been true.

I only noticed because I went looking for the failure mode deliberately, before it had a chance to bite anyone. I killed the assistant's listening process on purpose, left its credential completely untouched, and sent a test alert. It went out under the assistant's own name, formatted correctly, delivered successfully, with the dead process never going to read it. The old check looked like this:

```
if send_as(identity, message).ok:
    return "delivered"   # true even if nothing is listening
```

`send_as().ok` only tells you the API accepted the request. It says nothing about whether anything downstream will act on it. A credential working and a process being alive to use it are two separate facts, and I had built a check for exactly one of them.

The fix replaced the question entirely. Instead of asking whether authenticating as the identity succeeds, it asks whether the process behind that identity is currently alive, answered by a heartbeat that process publishes on its own, on a schedule, somewhere the sender can read it independent of whether sending itself would succeed:

```
if not heartbeat.fresh_within(identity, max_age=HEARTBEAT_WINDOW):
    fall_back()   # any read failure counts as "cannot act", no exceptions
else:
    send_as(identity, message)
```

If that heartbeat is missing, stale, or unreadable for any reason at all, the system falls back. It does not try to distinguish between the process being genuinely down, the heartbeat being unreachable, or some other read failure. Every one of those is the same outcome to the person waiting on the other end: they think someone is going to see this, and no one is.

That collapsing of causes into one outcome was the part I almost didn't do. My first instinct was to handle "definitely down" and "can't tell" as different cases, maybe log the second one as a warning and send anyway. But the person on the receiving end doesn't experience the difference. A message that silently lands nowhere because the assistant crashed costs them exactly as much as one that silently lands nowhere because a status check timed out. Refusing to send under that identity, and saying so out loud rather than failing quietly into it, is the only version of this that doesn't eventually cost someone a response they were counting on.

The fallback isn't silent either, which turned out to matter as much as the check itself. When the primary identity can't prove it's alive, the message still goes out, just under a different, clearly-labeled identity, with a line at the top saying why: that the usual assistant appears to be offline and this is a backup. If even that fails, it drops to a plain text message, also labeled. Every rung announces itself. An unannounced fallback is the version of this that goes wrong quietly for months, because the day the primary identity's token finally expires for real, everything downstream still looks fine. The label is what makes a silent failure visible instead of invisible.

I tested the whole chain by actually breaking the primary identity on purpose and confirming a real message arrived through the backup path, labeled correctly, rather than trusting the code to be right because it looked right. `send_as().ok` is still true the instant the process behind it dies; the heartbeat check is the only one of the two that would have caught it, and it is the one I did not write until I went looking for the gap on purpose.
