---
title: "The cockpit outage was just IIS taking a nap"
canonical: https://dxdev.com/blog/2026-05-13_cockpit-iis-nap/
datePublished: 2026-05-13
---
The site was hanging. Not erroring, not returning a 500, not timing out with a TCP reset. Just hanging. A request would go in and nothing would come back for thirty, forty, sixty seconds, and then eventually the response would arrive like the server had been thinking very hard about it.

My first instinct was application code. Something blocking on an external call, a database query sitting on a lock, an LLM request that had gone sideways. I pulled logs. I checked the process. Nothing looked obviously wrong.

The real cause was IIS's DefaultAppPool, which had quietly recycled itself after twenty minutes of no traffic.

## twenty minutes of silence and IIS goes to sleep

IIS has an idle-timeout setting on every application pool. The default is 20 minutes. If no requests come in for twenty minutes, IIS shuts down the worker process to save memory. The next request that arrives has to restart the whole process from scratch: load the application, initialize dependencies, warm up whatever caches the app keeps in memory. For a small side project serving sporadic traffic, this is fine. For an AI-facing app where request latency already has a floor of two or three seconds, a cold start that adds thirty seconds is indistinguishable from an outage.

The site was assigned to DefaultAppPool during initial setup. DefaultAppPool has idle-timeout on. Nobody touched it. The app had been sitting quiet long enough to get shut down, and then I came back to it and got thirty seconds of silence while IIS rebuilt the world.

## the fix was two settings and one new pool

I recovered the stalled pool from the IIS manager, confirmed the app came back responsive, then made two changes.

First: set `idleTimeout` to zero on the pool the cockpit was using. Zero means never shut down the worker process due to idle time. The pool stays warm as long as the app pool itself is running.

Second: moved the cockpit off DefaultAppPool and onto its own dedicated pool. DefaultAppPool is shared. Anything else running on the same IIS box that crashes, gets recycled, or hits its own resource limits brings down every app assigned to DefaultAppPool with it. A named pool means the cockpit has an isolated lifecycle. If something else goes sideways, it does not take the cockpit with it.

Both changes took about five minutes in the GUI. The site came up clean and stayed up.

## the fix that doesn't survive a rebuild

I sat with it for a minute after closing the IIS manager. The changes were live. The site was healthy. And I had not written any of this down anywhere a machine could read.

The `idleTimeout` setting lives in the IIS metabase, a machine-local XML configuration. If I rebuild the box, re-provision the server, or move to a new host, none of these settings come with it. They live as machine state, visible only to whoever next opens the IIS manager and thinks to look.

That is not a fix. That is a ritual. Somebody has to know to carve out the cockpit from DefaultAppPool and set `idleTimeout=0`, and right now that somebody is future-me, relying on memory.

So I opened the IIS setup script, the PowerShell file that provisions the box's initial configuration, and added the carve-out explicitly:

```powershell
New-WebAppPool -Name "cockpit-pool" -Force
Set-ItemProperty "IIS:\AppPools\cockpit-pool" processModel.idleTimeout ([TimeSpan]::FromMinutes(0))
Set-ItemProperty "IIS:\AppPools\cockpit-pool" recycling.periodicRestart.time ([TimeSpan]::FromMinutes(0))
```

Now when someone provisions a fresh box and runs the setup script, the cockpit gets its own named pool with idle-timeout off and periodic restart off, without anyone having to remember why.

## what the open question actually is

The incident note I captured ended with a question: what other the cockpit deployment assumptions are still living in machine state instead of the setup script?

I do not have a full answer yet. The IIS pool config was the visible one because it produced a hang. There are probably others. Log rotation. Scheduled tasks. Environment variable blocks that got set manually during a debugging session and never made it into a config file. Things that work right now because the machine is in the state it was in when somebody last touched it, not because the setup script would reproduce that state on a fresh box.

The honest answer is that I will probably find the next one the same way I found this one: after it fails.

## provisioning is where fixes live

The heroic restart, the manual GUI click, the one-time incantation at the command line: these are how you get a service back up. They are not how you keep a service up.

Every manual fix is encoding knowledge that isn't anywhere else. That knowledge belongs in the setup script, not in the memory of whoever happened to be on-call. The durable version of any infrastructure fix is the one a fresh machine would apply automatically.

Twenty minutes was all it took: one idle timeout on a shared pool turned a healthy app into thirty seconds of silence, and the only thing that stopped it was a GUI click that lives on one machine. The three lines in the setup script that create a dedicated pool and set `idleTimeout` to zero are the actual fix, because they are the only version of it a fresh box would ever apply. Until the other assumptions hiding in machine state get the same treatment, every one of them is a timer that has not run out yet.
