---
title: "When Markdown Defeated My Publish Gate"
canonical: https://dxdev.com/blog/2026-09-08_markdown-defeats-guardrail-regex/
datePublished: 2026-09-08
---
## Dead for five nights

At 6:58 AM on the sixth day, I found the nightly blog drafter had been dead since the previous week. Not crashed loudly, just quietly producing nothing, night after night, while the cron log reported success. When I finally traced it, there were four separate bugs stacked on top of each other. Three of them turned out to be things I'd assumed were broken and weren't. The fourth was a single word boundary in a regular expression, and it's the one that actually cost me time.

Once I cleared the pile, the drafter caught up and wrote 43 posts in one run. But the bug worth writing down is the one where the code was technically correct and still wrong.

## A gate that reads its own footers

Every draft that moves through the pipeline passes through a scrub step before anything gets written to disk: a list of regex rules that catch the company's internal shorthand and swap it for a public-safe phrase. One of those rules is deliberately boring:

```python
Rule("company token", r"(?i)\bacme\b", "prose", "the platform")
```

Case-insensitive, word-bounded, replace with "the platform." It had been running clean for months. The rule's job is to make sure the internal codename for the business never survives into a published post, and on every draft I'd ever manually checked, it worked.

What it missed was the drafter's own source material. The internal work-log entries that feed the drafter end with an attribution footer, and that footer writes the codename wrapped in markdown emphasis, like `_acme_`. That string carries the codename in it just as plainly as the word on its own. It should have tripped the rule. It didn't, and the survivor sailed straight through scrub and into the held-for-review queue, where the gate flagged it as unresolved and stopped the whole night's batch rather than publish something unscrubbed. Fail closed is the right default. It's also why the drafter looked dead instead of throwing an error I could grep for.

The reason `\bacme\b` doesn't match `_acme_` is that Python's regex `\b` treats underscore as a word character, same as a letter or digit. In `_acme_`, the boundary between the leading underscore and the "a" isn't a boundary at all as far as `\b` is concerned, both sides are "word" characters. The rule wasn't failing to run. It was failing to see the match in the first place.

## The fix that fixed the wrong bug

My first patch was the obvious one: stop treating underscore as a word character entirely, so it behaves like a space or a comma on either side of a match. I wrote it, ran it against the held drafts, and watched every `_acme_` survivor disappear. Good outcome, so I moved on and let it run.

Two nights later a different problem showed up: legitimate prose quoting real module paths, like `refresh_acme_status`, started getting flagged as company-token leaks. The codename now sat between two boundaries by the new definition, because underscore no longer counted as gluing it to its neighbors. The rule that was supposed to protect one specific word was now firing on the same three letters buried inside a snake_case identifier that had nothing to do with company disclosure. That's the mirror image of the original bug, and it cost me a re-review of every draft that had gone out in that two-night window to make sure nothing got mangled or wrongly held.

Underscore isn't uniformly a word character or uniformly a separator. It depends on where it sits. An underscore at the edge of a word, immediately before or after whitespace or punctuation, is doing emphasis. An underscore sitting between two letters is doing snake_case. The fix needed to tell those apart:

```python
_WORD_CHAR = "[0-9A-Za-z]"
_LEAD = rf"(?<!{_WORD_CHAR})(?<!{_WORD_CHAR}_)(?={_WORD_CHAR})"
_TRAIL = rf"(?<={_WORD_CHAR})(?!{_WORD_CHAR})(?!_{_WORD_CHAR})"
_BOUNDARY = rf"(?:{_TRAIL}|{_LEAD})"
```

`_LEAD` only counts as a boundary if there's no letter or digit before it, and no underscore-then-letter before that. `_TRAIL` mirrors it. Every rule pattern gets compiled once with literal `\b` swapped out for `_BOUNDARY`, so nothing upstream had to change, just the one compile site. That distinction, underscore-at-the-edge versus underscore-in-the-middle, is the entire fix. It's also the entire difference between a false negative and a false positive on the same three letters.

## What else is quietly assumed

The scrub rules were never wrong about what to catch. They were wrong about what a word boundary means once the text they're scanning can carry its own formatting. Plain text was never the actual input, it was just the only input I'd tested against. Anywhere else in the pipeline that reaches for `\b` and calls it done is making the same bet, and I won't know which one's wrong until it holds a batch for five nights and I go looking.
