The static scanner gave one Windows agent configuration a D, 57/100, after it found 24 issues and marked 14 of them high severity.

That is the kind of report that can turn into a bad policy decision fast. I was evaluating a config scanner because our agent setup had one security gap I could not already explain with a spend audit: prompts, hooks, permissions, and MCP configuration can all become an execution surface. The workstation had a public read-only knowledge-vault connector and a browser broker. The risk category was real, even if we had not yet had an incident.

So I ran the cheapest possible version first:

Terminal window
npx ecc-agentshield scan --path C:/Users/<username>/.claude

It was an offline static scan. I left off --opus, --injection, --sandbox, --taint, and --deep, which meant zero metered spend. The scan returned 0 critical findings, 14 highs, and a grade that said the configuration was in bad shape.

It was wrong about the bad shape. It was not wrong that one small piece was missing.

The report was describing a different machine

The first high finding said that CLAUDE.md was world-writable at mode 0o666. That would matter on a multi-user POSIX server, where another local account could edit an instruction file that an agent will later load.

This was a single-login Windows desktop. The operative permission system was Windows ACLs, not Unix mode bits. The scanner had converted a compatibility view of the file into an attack path, then graded that attack path as real without establishing that another user could actually write the file.

That distinction is not semantic cleanup. It changes the threat model.

A scanner finding is not a fact about a file. It is a conditional claim: if an attacker can reach this object through this permission model, then this configuration can produce this consequence. Here, the scanner got the file attribute and skipped the conditions.

The other highs followed the same pattern. I traced every one back to the source configuration instead of accepting the severity label.

Scanner signalWhat I checkedWhat it actually meant
CLAUDE.md marked world-writable at 0o666The Windows ACL context and who could log into the workstationA POSIX multi-user warning applied to a one-login Windows machine
A bootstrap hook and global instruction file flagged as attacker implantsThe hook text and the global configuration we wroteDeliberate bootstrap behavior: a PATH export, &>/dev/null, cargo install, and the <!-- rtk-instructions --> marker
Permission rules marked over-broadThe exact allow entriesA scoped python3 -c one-liner and find … -name .env*, which only lists names
A third-party check-config.sh script marked for possible leakageThe script body, its file reads, and its network behaviorIt trims a configuration key name and reads its own config. It does not exfiltrate data or make a network call.

The last one was especially useful as a diagnostic. The scanner had tripped on substrings in a clean third-party script from last30days. If I had stopped at the label, I would have treated a string match as evidence of data exfiltration. Reading the script took less time than deciding how to suppress the finding.

The bigger issue was coverage. The scanner discovered three skills because it saw plugin .md shims. It missed roughly 40 real skills under .claude/skills. That meant its D grade was not even a coherent score over the configuration that matters. It was a score over the subset it recognized, measured against a security model built around multi-user POSIX systems and proprietary skill conventions.

The project ships a false-positive-audit.md. That was an honest clue about the tool’s limits, but it also told me what running it weekly would look like. We would spend time re-litigating the same mismatched assumptions.

One finding survived the audit

The scan did find a cheap gap that was real on this machine. settings.json had an allow list and no deny list.

That does not become critical just because a scanner reported it. The existing allow rules still constrained what the agent could do. But a deny list is the correct second boundary for a small set of operations we never want available through an agent session by default: force-push, rm -rf, and raw secret reads.

I kept the fix narrow. The goal was not to make the configuration look secure to a generic rubric. The goal was to make a bad operation harder to approve accidentally, even when another permission rule changes later.

There was one piece of ordinary cleanup too. An old allow rule, Bash(find <retired-drive>/Projects/<legacy-project> …), still referenced a retired project drive even though active projects had moved to a different one. That was not a security exposure. It was stale configuration that could make a future review harder, so it was worth deleting while the file was open.

This is why I do not dismiss a noisy scan outright. A false-positive-heavy report can still be a useful forcing function if you separate the scanner’s evidence from its conclusion. The evidence gave me an inventory of configuration edges to inspect. The conclusion, the D, was not portable to this workstation.

Why this did not become a build gate

The obvious alternative was to run the scanner on every configuration change and treat a high finding as a build failure. I rejected that. A build gate is only useful when its failures are more trustworthy than the people who have to investigate them. Fourteen high findings that all collapse under a basic threat-model review would train us to ignore the next report.

I also did not enable the scanner’s optional --opus tri-agent pass to get a more elaborate verdict. The question in front of me was not whether a second model could write a better explanation of a POSIX permission warning. It was whether the file could be changed by a plausible attacker on this Windows machine, and whether the flagged commands could do what the report implied. That required config inspection and reading the referenced script.

A recurring scanner may make sense later if the environment changes. A multi-user host, broader execution permissions, untrusted plugins, or a larger set of writable MCP definitions would change the calculation. On this setup, the recurring cost would be review fatigue, not just runtime.

The result is deliberately unglamorous. We kept the current configuration, added the missing deny-list boundary, removed the stale path, and did not install another permanent security ceremony.

The D was not a diagnosis. It was a compact statement about somebody else’s attacker, operating system, and configuration conventions. Once I unpacked those assumptions, every high finding was noise. The one useful result was a small control that should have been there anyway.

That is the bar I want for agent security tooling. Show me the path from a real attacker to a real consequence on the machine I actually run. A letter grade can start the audit. It cannot finish it.