---
title: "Three Safety Configs, and Nobody Had Compared Them"
canonical: https://dxdev.com/blog/2026-09-20_three-safety-configs-nobody-checked/
datePublished: 2026-09-20
---
The diff command returned zero lines of output. Two `settings.json` files, one for my main work profile and one for the profile my business partner and I both run client work out of, five enforced hooks apiece, byte-for-byte identical. That was the expected result, and it's the result that almost made me close the ticket.

The audit started as something else entirely: a pass over roughly 67 loose skill files that had accumulated in `.claude/skills/` with no index, no dedup, and a hand-maintained README that documented 12 of them and hadn't been touched in over two weeks. While a subagent walked the skill tree, I turned to the second half of the same question: hooks. Forty-one hook script files sat in one shared directory, wired into whatever `settings.json` each work profile pointed at. I knew of two profiles. I compared them, got the zero-diff result above, and started writing up "hooks are consistent across seats" as a finding.

Before I wrote the last line, I ran one more grep, this time for every `settings.json` on the machine instead of just the two I already knew about. It came back with three. The third wasn't in my mental model at all. I found it because a hook script's own docstring referenced a config path I didn't recognize, pointing at a profile I'd never accounted for: the one that runs this blog's own posting and automation work. It had been running the whole time, quietly configured, never compared to anything, because nobody had ever framed the question as "three profiles" instead of "two."

Once I diffed all three instead of two, the picture changed. The third profile was missing three hooks that both of the other two enforce:

- `estimate_gate`, which blocks a dispatch from starting without a duration estimate on record
- `reply_shape_gate`, which checks that a generated reply matches the required output shape before it ships
- `ask_shape_gate`, which does the same check on the question-asking side of an agent exchange

And it was carrying two hooks the other two didn't have at all: `dispatch_budget`, a spend guard on metered API calls, and `hook_block_direct_browser_drive`, which forces browser automation through a supervised path instead of a direct driver call.

None of this was a deliberate decision. There's no comment anywhere saying "this seat doesn't need reply_shape_gate." The two identical profiles got that way because one `settings.json` was copied into a second directory at some point and never diverged. The third profile was built separately, on its own timeline, for a narrower purpose, and picked up whatever hooks made sense for that purpose at the time. Nobody subtracted the enforcement hooks on purpose. They just were never added, because the profile was never compared against the other two, because until this audit there was no artifact that listed all three side by side.

The part that actually stopped me was `reply_shape_gate`. That's the hook that catches a generated reply before it ships if it doesn't match the required shape, no leading "Great question," no meta-commentary about the draft, no sign-off tacked onto the end. It's a small, boring, easy-to-forget guard. And it was absent on exactly the one profile whose job is generating posts for public consumption. Not because anyone decided that profile didn't need it. Because the audit that would have caught the gap hadn't been run, and the seat itself had surfaced by accident, from a docstring, mid-investigation.

Two profiles being identical wasn't evidence the system was well-governed. It was evidence that nobody had looked hard enough to find the third. A diff on the pair I already knew about would have told me everything was fine, and everything was fine, for the two I asked about. The finding wasn't in the comparison. It was in noticing there were more things to compare than I'd assumed, and going and counting them instead of trusting the count I walked in with.

The fix is a generated inventory that enumerates every seat's hook set from the filesystem itself, rebuilt on every audit instead of assumed from memory, so "how many profiles do we actually have" stops being a question anyone has to remember the answer to.
