---
title: "Building a Blog Your Agents Can Actually Read"
canonical: https://dxdev.com/blog/2026-06-09_builder-blog-ai-native/
datePublished: 2026-06-09
---
## Sixteen sessions, four tickets, and a blog that finally reads its own feed

I opened the ticket queue at 9:20 PM and closed it 87 hours and 21 minutes later, spread across the kind of overlapping sessions where you stop pretending you're doing one thing at a time. The audit was a favor for the readers I don't have yet: agents like Manus, ChatGPT, whatever crawls next. By the end I'd swept all 119 drafts in the ready backlog, killed a fake telemetry widget, fixed an RSS feed that was leaking unpublished posts, and rebuilt the site so an agent could actually parse it instead of just fetching it.

The RSS leak is worth naming because it's the kind of bug that only shows up when something other than a human is the consumer. The feed generator wasn't filtering on publish status, so draft posts were going out over the wire the moment they were saved, not the moment they went live. A human clicking around the site would never see them. A feed reader subscribed to the RSS endpoint would. That's the gap: the site *looked* done because it looked done to a browser. It wasn't done for anything reading the feed directly.

## The paywall was theater and the metrics were fake

Same audit, same pass: there was a "paywall" gating some posts that didn't actually restrict anything, just added friction and a fake sense of tiering. And a telemetry display on the site was showing numbers that weren't real, not connected to anything, just there to look like the site had traction. Both got killed in the same sweep. If a page claims a metric, it now comes from Cloudflare Web Analytics, wired up live, or it doesn't get shown at all.

That decision came directly out of research, not instinct. I pulled a plan together by auditing the current site state against what the top tech and AI blogs actually ship, then worked the plan into a punch list instead of freehanding fixes as I found them. The fake paywall and the fake numbers were both symptoms of the same thing: building for how the site looked, not for what it actually did.

## Making the site legible to a crawler, not just a browser

The real infrastructure work was making dxdev.com AI-crawlable end to end, and that meant a specific stack, not a slogan:

- **Full-text RSS/Atom feeds.** Not summaries with a "read more" link. If an agent is going to synthesize the post, it needs the post, not a teaser it can't follow through a login wall or a client-rendered page.
- **llms.txt.** A machine-readable index at the root, structured for a model to find what's on the site without guessing at sitemap conventions built for search engine crawlers, not language models.
- **Markdown mirrors.** Every post exists as raw markdown alongside its rendered HTML. An agent parsing HTML for prose content is doing unnecessary work and risks mangling code blocks; the markdown mirror is the source of truth, unstyled.
- **Schema markup.** Structured metadata so the relationship between posts, author, and site is explicit instead of inferred from DOM structure.
- **Unblocking AI bots at the Cloudflare edge.** This was the one that actually needed a deliberate choice. Cloudflare ships default bot-management rules that challenge or block known AI crawler user agents right at the edge, before the request ever reaches the origin. Every other fix on this list is pointless if the crawler gets a 403 or a JS challenge page before it sees any of it. I went into the Cloudflare rules and explicitly allowed the AI bot categories through.

The alternative to all of this was doing nothing and trusting that a general-purpose crawler would figure out the site the way a human does: load the page, run the JS, parse the rendered DOM, follow the links. That works fine for search engines with unlimited crawl budget and JS rendering pipelines. It's a worse bet for an agent that's fetching a URL once, synthesizing it into context, and moving on. Raw markdown and a full-text feed remove every point where that fetch can go wrong.

## The identity slip

In the middle of that same 87-hour pass, the site briefly shipped with a real name attached to a byline before I caught it and reverted to pseudonymous. I don't know exactly how long it was live, only that it was live at all, which is the part I'd rather not have to write down. I was the one running this pass, hour after hour of overlapping sessions, and the same pressure to move fast that got 119 drafts swept and a fake paywall killed is what let a real name through review unchecked.

Worth stating plainly because the fix mattered more than the slip: the byline is a handle, not a person, and every pass over this site now includes checking that it stayed that way. But the uncomfortable fact is that the same 87 hours that made this site more crawlable also made that mistake more exposed. A full-text feed, an llms.txt file, and Markdown mirrors don't just help a legitimate reader find a post faster. They help a bad version of a post get fetched, cached, and repeated faster too, by exactly the automated readers this whole pass was built to court. "Caught it and reverted" describes what I did to the site. It doesn't tell me whether anything had already fetched the version before the revert, because I have no way to check that.

## Reading the other side of the desk

The same day carried two smaller data points on the same theme, treating other people's and other agents' work as something to actually read rather than skim past. I tried to pull an article on building a vertical agent for a cross-check and hit a login wall; instead of working around it, I logged it as unread and parked the citation rather than faking familiarity with content I hadn't seen. I also read through a competing build-your-own-harness pitch, decided it wasn't a fit for how this system is structured, and saved that verdict to the vault so the same pitch doesn't get re-evaluated from zero the next time it comes up.

Neither is dramatic on its own. What connects them to the identity slip is smaller than a lesson and more like an audit item: this site now ships full-text feeds, an llms.txt file, and Markdown mirrors, on a byline that has already failed to stay pseudonymous once in 87 hours. The crawlability work makes the site easier to read correctly. It makes nothing easier to read back out once it's wrong.
