---
title: "The 5-Minute Poll That Made Orchestration Feel Broken"
canonical: https://dxdev.com/blog/2026-09-26_orchestration-ack-latency/
datePublished: 2026-07-30
---
The evening my partner tested my agent with a money-move ask in #agents, the first thing I noticed was the wait. The request was handled correctly. But nothing visible happened for up to 5 minutes, and in a shared channel a silent agent looks broken.

## What the poll was doing

The agent woke on a 5-minute poll, fetched whatever had arrived, and worked through it. Throughput was never the problem. A message that landed one second after a poll waited nearly the full interval before anything acknowledged it. I had tuned the work and left the front door alone, because the work was the part I could measure.

Nobody who saw that silence could tell "thinking" from "dead." My partner was probing on purpose, so he was watching for exactly that.

## The first fix was the wrong one

The probe came in and I went straight to the money-move problem, since a request to move money is the case where an agent must not be helpful. I built quarantine: one reply, then silence. It covers the person and their agent, across all channels. It can be lifted only from a room they cannot reach. I also added a Discord DM review loop, so I hear about the probe and answer in my own words instead of the agent improvising.

That worked as designed, and it exposed the second problem. Every wake saw one message and no thread. In a DM review loop the whole point is a conversation, and each wake started from zero. I answered a fragment, then answered the next fragment, and the agent had no memory that the first one existed. I had wrapped a safety mechanism around a delivery mechanism that dropped context, and I had to fix the delivery underneath it after shipping the safety on top. The amnesia was a symptom of the same design as the latency: a poll hands over a snapshot, and a snapshot has no history.

## Replacing the poll with a websocket daemon

The replacement is a websocket daemon. It holds a connection open, so a message arrives as an event the moment it is sent, and the acknowledgment goes out in under a second instead of on the next 5-minute tick. The agent still takes as long as it takes to do the work. What changed is that the sender sees "received" almost immediately, and the work behind it is allowed to be slow.

Because the daemon owns the connection, it can also keep the thread. The amnesia fix was to stop treating each wake as a fresh start and hand the agent the conversation, not a single message.

## Why the ack matters more than the answer

The latency is worth measuring separately from the work. A 5-minute poll and a fast acknowledgment can sit in front of the same stretch of real work, and the two feel completely different. The first feels like an outage. The second feels like a colleague who said "on it."

The same day, a protocol review I ran on the agent-to-agent setup came back with a verdict I keep thinking about: it was not safe to leave unattended as written, because the mechanical limits controlled message volume while the important decisions, what is authorized, were not covered. That is the same shape as my ordering mistake. I had capped and timed the plumbing and never looked at what the plumbing was allowed to do or how it looked from outside.

## What I would check first next time

1. Time from message sent to first visible sign of life, measured from the sender's side. Not time to finish.
2. Whether a wake sees a thread or a single message. If it is a single message, every multi-turn flow is quietly broken.
3. Whether any safety layer sits on top of a transport that cannot support it.

I closed the private agent-to-agent channel the same night. The acknowledgment path is now a socket and not a schedule, and the thing my partner poked at answers before he has looked away from the screen.
