---
title: "I stopped babysitting the browser, then had to build the part that was missing"
canonical: https://dxdev.com/blog/2026-08-22_multi-session-browser-isolation/
datePublished: 2026-06-26
---
At 10 concurrent agents, the browser stopped being a tool and became a scheduling problem.

I had already made one agent's browser work. I swapped the control server, opened a real Chrome window, and confirmed that a fresh session could connect. That solved the immediate complaint.

Then I ran it the way we actually work. Several sessions were active across repositories. One agent took over another agent's page. A window appeared offscreen. A session said it could not open a browser. Worst of all, a launcher could reach my personal Chrome profile instead of the controllable browser the agent was supposed to use.

The original fix was a configuration change. The real problem was browser infrastructure.

## One Chrome was doing too many jobs

The first design looked efficient. One Chrome instance listened on port 9222. A broker accepted connections from sessions and gave each one a separate Chrome DevTools Protocol context. An older 25-slot pool existed, with ports `9300 + slot` and one profile directory per slot, but the shared path was active.

That saved processes. It did not produce isolation I could trust.

The diagnostic path mattered. The broker had one child process, one cached initialization result, one writer queue, and a tool table scoped to the active page. When two sessions became busy at once, the apparent agent conflict was shared browser state colliding inside a single-upstream component.

A code review exposed a second problem. `acquire()` had no cross-process lock. Two sessions could race and receive the same slot. That was the clobbering root. It was not a prompt problem, and it was not something I could solve by asking agents to be more careful about ownership.

The quiet failure was worse. If session identity resolution missed, the old wrapper still started. The broker invented an ephemeral identity and put the session into the shared default context. A failed lookup could silently disable the isolation the system existed to provide.

I changed that behavior to a hard failure. If the wrapper cannot resolve a session identifier, it exits with `rc=2` after a roughly 10-second wait. A broken browser is visible. A browser that looks isolated but is not is worse.

## I did not extend the broker

The tempting path was to make the shared broker handle multiple upstream browser processes. That lost once I inspected the code. It would mean rebuilding the component with the most shared state, then probably deleting it later.

I kept the old shared path as a rollback switch and built the new path beside it. The wrapper now acquires a slot, ensures the slot's Chrome process is alive, connects to that slot's debugging port, and starts its own DevTools control process. There is no central process deciding whose page is active.

Each slot gets its own persistent profile directory, port, and Chrome process. Allocation is protected by an interprocess lock. If the port is dead, the launcher clears stale singleton files and starts the browser again. The profile remains intact so cookies, local storage, and IndexedDB survive a restart. Only cache directories are cleared on release.

I considered resetting the profile every time. That would have been simpler to reason about, but it would turn normal agent work into a login treadmill. An exported storage snapshot also lost, because it cannot restore a full profile and session storage is separate.

The launcher is pinned to Chrome for Testing, not the system browser. It fails loudly if that binary is missing. It persists the Chrome PID, requires a profile directory, and strips remembered window placement before launch. The old launcher used the system browser, dropped the child PID immediately, and missed the workaround for invisible Session-0 windows.

## A CDP context is not a window

The shared design could create logical browser contexts. It could never make ten agents visible at once. I need to see what an agent is doing while it is doing it, especially when work crosses staff tools, customer-facing checks, and issue tracking.

The fleet uses real headful windows on a dedicated monitor. The launcher reads monitor geometry, accounts for DPI, and places the first four windows in a 2x2 grid. It moves a window without activating it. When I want to inspect a session, the control command resolves the saved launch PID to its window handle, restores it if needed, and brings that one window forward.

That fixed two complaints. Agents no longer hid in offscreen windows, and they no longer stole focus every time a session touched the browser. Visibility is the default. Foreground is an explicit request.

I rejected a screenshot wall as the primary interface. A DevTools screencast may be useful later as a dashboard, but it cannot replace a browser window that I can inspect and interact with. I also rejected headless overflow. If the system says it has a visible browser fleet, the eleventh agent does not get an invisible exception.

## Login state has an owner

Persistent profiles create their own risk. If an arbitrary session can land in an arbitrary old profile, it can inherit someone else's login.

The pool assigns a login target, not just a slot number. The wrapper derives a target from an explicit override when present, then the repository root, then a default. An affinity ledger maps that target back to the same slot on later sessions. The first use requires a manual login. After that, the repository gets durable browser state instead of asking us to log in again for every session.

If a slot must be reassigned, the system wipes authentication first. That prevents cookie bleed. With 25 slots and roughly 12 expected targets, reassignment should be rare, but the safeguard cannot depend on that expectation.

## The cap is part of correctness

A Chrome process costs roughly 300 to 600 MB. I set the concurrency cap at 10. When the fleet is full, a new browser request waits and reports why. An operator can explicitly reclaim the oldest idle live slot. The system never quietly opens headless Chrome to get around the cap.

I also made the wrapper lazy. A new agent session does not acquire a slot or launch a window during its handshake. It does that only on the first real `tools/call`. This removed the failure mode where idle sessions filled the pool before doing any browser work.

At first I considered a periodic zombie reaper. The old lease only had a 30-minute timestamp, so a reaper could have killed a live session that was simply quiet. I deferred it until the pool had heartbeat and activity tracking. Cleanup is safe only after the system can distinguish idle from alive.

The final configuration is small:

```text
AGENT_BROWSER_BROKERLESS=1
AGENT_CHROME_CONCURRENCY_CAP=10
```

The work behind those two lines is not small. One agent with one browser is a configuration problem. Ten concurrent agents with separate state, persistent logins, visible windows, and a reliable failure boundary is a pool, an identity system, a process supervisor, and a capacity policy.

I stopped babysitting the browser when I stopped treating it as a disposable window.
