---
title: "I Wrote Six Rules, Then Graded My Own Published Posts Against Them"
canonical: https://dxdev.com/blog/2026-09-22_six-rules-for-a-blog-an-ai-writes/
datePublished: 2026-09-22
---
## The score that lied

quality6: 6. avg: 84. Every one of six gates green. I pulled up the post underneath that score to spot-check it and it read like a press release: no ticket counts, no field names, no numbers at all. Somewhere in the pipeline, a rule meant to protect identity had eaten the whole post instead.

That was the moment I stopped trusting my own grading script, and it's worth walking through how it got built and how it broke.

## From vibes to a rubric

The first pass at auditing the blog wasn't a script, it was a single long session: point an agent at all 72 published posts at the time, ask for a duplicate map, quality tiers (Flagship, Solid, Weak, Cut), and a recommended launch set. It worked. It caught real overlap, flagged one post (the April diagnostic writeup) as an unpublished duplicate to treat as cut, and named the sameness pattern that shows up after ten posts in a row.

But it was one read, by one agent, on one day. Ask again next week and you'd get a different Weak/Solid split, because there was no written standard behind the tiers, just judgment. That's fine for a one-time audit. It's not fine as a permanent quality bar for a corpus that had grown to 187 posts and a drafts queue behind it.

So I wrote the standard down. Six rules, each phrased as something a script could actually check: no em dashes, at most one "X, not Y" construction per post, anonymize by substituting an equivalent specific rather than deleting it, open on a hard number or symptom, no templated headings, no closing line that just points at a companion post. Six gates, pass or fail, per post.

## Wiring the gates into the pipeline

The grading script runs every published post and every queued draft through repeated judge passes and writes one record per slug: `avg` (0-100, averaged across passes), `spread` (how much the passes disagreed), `gates_green` (how many of the six rules it cleared), `reframe` (keep or cut), and a one-line `suggestion`. All of it lands in a single `ranked.json`. A second, much smaller script does the actual triage: read `ranked.json`, sort by `avg` descending, drop anything on a manually maintained drop list or already-done list, and take the top 32 as `editorialTargets`, the posts worth a real second look this cycle.

That second script is maybe fifteen lines. It's the first one, the gates themselves, that had the bug.

## Where the anonymize gate went vacuous

The anonymize rule started out as "strip anything that could identify a person, company, or customer." Under that wording the gate was satisfied by deletion. A post that used to describe a specific traffic pattern in numbers could have every number stripped, and the gate would score it green, because there was nothing left in the post that could possibly leak anything. Vacuously true. Zero signal, six gates green.

I ran the full corpus through that version, got a `ranked.json` back with a healthy chunk of `gates_green: 6` posts, and fed the top scorers straight into the editorial-targets list without rereading them. That's the cost: a batch of "keep" recommendations that were actually gutted, and I only found out by manually rereading a sample afterward. One of them had a specific number in its own URL slug and never used that number once in the body. The gate hadn't protected anything, it had quietly deleted the article and called it compliant.

The fix was to change what the rule rewarded. Not "strip anything sensitive" but "substitute, not delete": replace the identifying noun with an equivalent generic one and keep every number, threshold, price, error string, command, and field name exactly as it was. A 63-team league instead of a named client. A tracker id instead of a real one. The count and the dollar figure stay untouched either way. I rewrote the gate check to specifically fail a post that anonymized a proper noun but also dropped a number in the same paragraph, on the theory that real anonymization changes nouns, not digits.

Re-running the corpus against the fixed gate demoted a real slice of the previous "keep" batch back down, correctly this time, because they'd been passing on emptiness rather than on actual compliance.

## What the rubric is actually for

A rule a script can satisfy by deleting the thing it's supposed to protect isn't a rule, it's a hole with a green checkmark on it. The useful version of self-grading isn't "did this pass," it's "can this gate be won by making the post worse," and that second question only shows up if you go read the output instead of trusting the aggregate number. `quality6: 6` is a summary. It is not a substitute for opening the file.

The grading pass isn't a one-time audit anymore. It runs before anything ships, against the full 187-plus queue, and the six gates are the actual constitution the drafts get held to, not a description of good writing I'm trusting an agent to intuit fresh every time.
