---
title: "I built a weekly audit for my AI mistakes"
canonical: https://dxdev.com/blog/2026-05-21_weekly-audit-ai-mistakes/
datePublished: 2026-05-21
---
The same correction appeared three times in one week: "wait, I thought we already committed that." Ground state before action is a rule in my MEMORY.md, filed under `feedback_ground_state_before_action.md`. It has been there since April. The agent reads it on every session open, and it still fails on it roughly once every three days. I caught each lapse individually, treated it as a one-off, and moved on.

My mental inventory of where Claude sessions went wrong was a collection of vibes: something about premature close, something about scope conflation, something about manifest drift, none of them sorted by frequency, none confirmed to be in the memory file yet, none distinguishable from a one-session fluke. Corrections I made on Monday sometimes became new MEMORY.md rules and sometimes vanished back into the transcript. I couldn't tell which without scrolling through the session history, and after the third time doing that in a week, I stopped. The corrective loop was running entirely in my head, and it was lossy.

So I built a weekly retro script.

## The audit that runs whether I remember to or not

The retro script lives in my ops repo. Every Sunday at 3 AM, a Windows scheduled task calls it. The script walks every transcript file under `C:/Users/<username>/.claude/projects/`, greps for correction and frustration markers, and clusters the hits by session and day. The markers are phrases I actually type when a session goes wrong: "wait," "no," "that's wrong," "I thought we," "hold on." Not every instance is a genuine correction, so the script uses surrounding-turn context to filter routine disagreements from actual redirects. Then it cross-references each cluster against MEMORY.md. Any rule that matches a correction cluster gets flagged as fired-but-failed: the rule exists, the agent read it, the session still produced the correction. Rules with no matching clusters get marked fired-ok, and rules that never appeared in any session this week get marked not-fired, which is its own signal.

The scheduling decision matters more than the script does. I had tried "I'll review my AI mistakes at the end of the week" twice before. Both attempts died before week three. When the sprint fills, the optional review is the first thing to go. Sunday at 3 AM has no opinions about the sprint, and I can read the output over coffee without having to remember to generate it.

## Why structure matters more than coverage

A flat list of complaints gets overwritten in memory by whatever happened in the most recent session. The design that actually works: stable section names, consistent rule slugs, the output landing at `vault/04_Sessions/_retro/` where each weekly run diffs against the prior one. If `feedback_ground_state_before_action` appeared as fired-but-failed in week 3 and week 4, both rows show up. If a rule went quiet in week 5, the blank row is visible rather than silently absent. The trend is the evidence.

My MEMORY.md has roughly 80 rules at this point. Most are `feedback_*.md` files with slugs like `feedback_no_close_before_work_done` or `feedback_swarm_scan_proper`. The retro reads each slug, finds whether any transcript hit matches it semantically, and marks it fired-ok, fired-failed, or not-fired. Duplicate rules surface as overlapping clusters. Rules that haven't fired in three months appear quiet in the report, which sometimes means the agent internalized them and sometimes means I stopped doing the work that would trigger them. You can't tell which without the trend line, and the trend line only exists if the section names stay identical run to run.

Getting the output shape stable enough to diff was the harder part of the build. The grepping took an afternoon.

## What the first full run found

The first seven-day run came back with twelve rule clusters and one finding I had not been able to name before seeing it. Session-footer cross-contamination, ranked first by hit count.

The footer is the auto-generated block at the end of a session manifest: session ID, area, epic, task, current status. Mine was leaking context sideways. A session opened for one ticket would inherit stale content from a prior session's footer, because the manifest builder was appending instead of replacing. The agent would start confident it was working on the previous ticket because that was the last thing stamped in the footer layer, even when the actual open session had nothing to do with it. Opening exchanges felt off-track for days. I kept attributing it to sloppy session openers, never landing on a concrete cause.

The retro found it because "I thought we" appeared 7 times across 4 sessions in one day, all in the first three turns of each session. The rule `feedback_ground_state_before_action` was flagged as fired-but-failed on 5 of those. But that rule addresses the agent skipping a state check before acting. These sessions had a different failure mode: the agent started with wrong state already baked into the context, a footer from a prior session stamping the wrong ticket ID into the new session's ground truth. The rule was right to fire, but the correction it pointed toward was wrong. Fixing the rule further would have changed nothing. The bug was in the footer writer, one layer below where the rule lives.

I patched the manifest builder the same morning: two lines that wipe the footer section before rewriting it instead of appending to it. The sessions that opened the next day were clean.

## What stays invisible without a diff

Most AI workflow frustrations live as vibes because vibes don't require evidence to sustain themselves. "Claude keeps doing this thing" is enough to feel annoyed. It is not enough to act on systematically, or to verify that the last fix actually worked.

The same decay pattern shows up in infrastructure. The 80-rule MEMORY.md looks healthy. You added 12 new rules this month. You can't see that the agent ignored 9 of the oldest ones without a run that compares this week against last. The ignored rules are there in the fired-but-failed column, waiting for a structured scan to make them visible as a pattern rather than noise.

The retro also distinguishes between a rule failing occasionally and a rule failing in a specific context. An occasional fired-but-failed hit might be a hard session. Five hits in one day, all in opening turns, all against the same rule, is a signal with a direction. That directional quality is only visible when you can compare it against prior weeks where the same rule was performing fine in those same positions. The footer contamination didn't need a smarter rule or a more attentive operator. It needed `feedback_ground_state_before_action` to appear in the fired-but-failed column specifically in opening-turn corrections, two runs in a row, which is how a pattern that looked like noise becomes a two-line fix on a Sunday morning.
