On June 13, the RSS feed still had a draft leak, and the ready backlog contained 119 drafts.
That was the point where I stopped treating crawlability as a nice publishing extra. We were bringing dxdev.com to a launch baseline, and the job started as an editorial cleanup. It quickly turned into a data integrity problem.
A blog about agents that only works for a human with a browser is making a strange promise. We write about agents, automation, long running systems, and the work that breaks when the happy path ends. An agent should be able to reach the same primary material a reader can reach, without guessing which rendered fragments are the article, which pages are drafts, or whether the site is putting on a fake product show.
The draft leak made the failure concrete. A feed is not merely a distribution channel. It is a public read API. If it includes unpublished work, it is wrong. If it contains only a teaser, it forces every downstream reader to reconstruct the article from HTML. If the visible site says one thing and the feed exposes another, we have created two contradictory records.
The audit started with the parts that were pretending
The first pass through the site found three separate problems that looked unrelated from the front end: fake telemetry, a theater paywall, and an RSS feed that could leak drafts.
The fake telemetry and paywall were not harmless decoration. They told visitors there was measurement and access control where neither existed. An automated reader cannot reliably infer that a counter is invented or a barrier is performative.
We removed both instead of decorating them more convincingly. The rule was simple: publish the capabilities we actually operate, and no others. We then put real measurement in place with Cloudflare Web Analytics.
The feed bug had a different shape. It was not a credibility issue. It was a boundary failure. Draft status had not been enforced at the point where the feed was built. The repair was to make the publishing state govern the feed, then sweep the 119 ready drafts so we were not carrying a pile of hidden exceptions into launch.
That diagnosis changed the rest of the work. The question was not, “How do we make a blog look ready for AI?” The question was, “What public interfaces expose the canonical article cleanly, completely, and truthfully?”
One article, several honest surfaces
We shipped four pieces of reading infrastructure together: full text feeds, an llms.txt file, Markdown mirrors, and structured schema.
They overlap on purpose. Each solves a different retrieval failure.
A full text feed gives a reader the whole post in an ordered stream. A summary feed lost here because the missing body is the point. The commands, failure modes, thresholds, and sequence of events live there.
llms.txt gives an agent a compact map of the site. It is not magic and it is not a ranking trick. It is a declaration that says, in a form a text system can consume cheaply, what this publication is and where its substantial material is. A crawler can still read navigation, cards, and rendered HTML. It should not have to infer the primary corpus from those things.
The Markdown mirrors are the durable reading surface. HTML is the presentation layer. Markdown preserves headings, links, code, and paragraph boundaries without pulling layout, scripts, newsletter modules, and theme chrome into the context window.
Schema gives machines an explicit description of the page and its relationship to the rest of the site. It is not a replacement for readable content. The feed, llms.txt, mirrors, and schema had to agree on what an article was.
The edge was the part that could silently erase all of it
I found that out after I had already deployed. The build check had passed, with 41 published posts in every surface and no draft slugs, and I deployed on the strength of it, treating “AI-crawlable” as settled. It only proved what the build emitted. A deploy check that requested the live site as GPTBot, ClaudeBot and PerplexityBot got a 403 on every request, the Markdown mirrors included. Cloudflare had also injected its own managed robots.txt block above ours, disallowing the same bots.
The public documents were only half the system. Cloudflare was sitting in front of the site, which meant a bot policy at the edge could make every crawlability improvement irrelevant.
That was the diagnostic path I wanted recorded. A missing agent citation can look like a content problem. It can look like bad schema, a weak feed, missing Markdown, or an indexing delay. But if the request is stopped before it reaches the site, none of those explanations matter. The markup can be perfect and still be invisible.
We unblocked AI bots at the Cloudflare edge after the text surfaces were in place. The order matters as a design discipline even when the changes ship together. First make the document canonical. Then expose it through interfaces that preserve its full body. Then make sure the network layer allows the intended readers to reach those interfaces.
I considered treating this as a one file change. Add llms.txt, declare victory, move on. That would have been theater in a new form. A single discovery file does not fix a feed that publishes drafts. It does not make an article body available as clean text. It does not resolve disagreement between the visible page and structured metadata. It certainly does not override a blocking edge rule.
Crawlability is part of the build
We build agent systems against repositories, logs, manifests, and operational records because a system cannot audit what it cannot read. A public technical blog deserves the same treatment.
The useful standard is whether an agent can find the canonical text, identify it as published, retrieve all of it without presentation noise, and trace it to a stable public source. That is infrastructure, not content marketing.
The RSS leak made the stakes obvious. The 119 draft sweep forced us to define a real publishing boundary. Now, when we publish a build log, we are giving agents a record they can actually read.