The oldest SSH tab dropped first whenever an agent spent several minutes thinking.

At first I blamed the iPhone.

That was a reasonable diagnosis. I was running multiple SSH sessions from a phone, one tab would sit open while an agent worked, and mobile operating systems are not famous for protecting background network connections. The failure looked like backgrounding. I would return to the oldest tab, type a command, and get nothing useful back. The newer tabs were still alive. If I moved between apps, the oldest session was usually the one I had to reopen.

The explanation was plausible enough that I nearly stopped there. That would have been a mistake.

The ordering was the clue

The failure was not random. It had a strict order.

The oldest connection died first. Newer sessions survived. A tab that was actively printing output survived. A tab that had gone quiet while the agent reasoned was the one that disappeared.

That distinction mattered. If iOS backgrounding were the whole story, I expected the visible app state to be the main variable. A session could fail after the phone moved it out of the foreground, regardless of when the SSH connection started. What I actually had was closer to an idle-timeout curve. The connection with the longest period of silence was consistently the first one removed.

The agent’s thinking phase made the pattern easier to miss. A long reasoning step looks like work, but it is quiet from the terminal’s perspective. No prompt is being typed. No command is producing output. The session is open, the agent is busy, and the network path sees an old TCP connection that has stopped sending packets.

That is a very different failure mode from a mobile SSH client being suspended.

I checked the wrong layer first

I began with the client because that was the part I could see. The phone was in my hand. The tab was on screen. The broken connection appeared after the SSH app had spent time in the background. That is how a plausible story gets ahead of the evidence.

I went to the logs instead.

The sshd configuration showed that server keepalives were disabled. There was no periodic server-originated traffic to ask whether an apparently idle client still had a usable path. The Tailscale logs pointed in the same direction. The connection was not being selected at random. The quiet path was aging out.

The actual chain was straightforward:

  1. I opened several SSH sessions through Tailscale.
  2. One session went quiet while an agent thought for minutes.
  3. The NAT state for that quiet TCP flow aged out.
  4. sshd had no server keepalive configured to generate traffic before that happened.
  5. The next interaction exposed a connection that had already been removed somewhere between client and server.

The important part is step 3. TCP does not get a vote when a NAT device decides its state entry is stale. From each endpoint’s point of view, the session can look open until one side tries to send something. That is why the tab did not always announce its death when the timeout happened. It often looked fine until I came back and used it.

The oldest tab was not unlucky. It had simply been silent longer than every other session.

The fix belonged on the server

There were several ways to respond.

I could have changed the phone app’s keepalive setting. That might have worked for that one client, but it would have made the behavior dependent on every SSH client I use. I could also have added a local loop that typed harmless commands into idle terminals. That would have been more machinery to maintain, and it would blur the difference between an interactive session and a healthy network path.

I could have treated this as an iOS limitation and accepted periodic reconnects. That lost because the logs had already contradicted the diagnosis. The same quiet-connection failure could affect another client on another device if it used the same path and left the session idle.

The durable control point was sshd. I enabled a nonzero ClientAliveInterval in its configuration.

The exact interval is a policy choice. It needs to be shorter than the relevant idle timeout, without turning every dormant terminal into needless traffic. The mechanism is what mattered here: sshd now sends periodic client-alive requests during an otherwise silent session. A quiet terminal stops being indistinguishable from an abandoned TCP flow.

I did not need to redesign the agent workflow, replace the phone client, or build a reconnect supervisor. One server setting restored the missing signal.

What the logs changed

Before I read them, I had a neat explanation: the phone backgrounded an app, so the app lost its connection. It fit the device in front of me. It fit the timing well enough. It also encouraged a client-side fix.

After I read them, the pattern was much less flattering to my first guess. The phone was incidental. The decisive facts were the disabled server keepalive, the Tailscale path, the NAT timeout, and the age of the quiet connection. The failure sequence explained why the oldest tab died first, why a busy terminal stayed alive, and why the disconnect surfaced only after the agent finished thinking.

This is the kind of operational bug that rewards a boring question: what is the system actually doing while it appears to be idle?

In this case, the answer was not “nothing is wrong until I touch the tab.” The connection had already died. I just had not asked it a question yet.

The fix is one line. The diagnostic is the part I want to keep: when a supposedly random disconnect has an ordering, treat the ordering as evidence. It is usually the system telling you which clock is running.