---
title: "I Stopped Guessing Categories and Read Them Off My Own Product"
canonical: https://dxdev.com/blog/2026-09-07_taxonomy-read-off-the-live-nav/
datePublished: 2026-07-27
---
## Measure before you architect

The alert wasn't an alert. It was a number: 82 tickets, 382 comments, and nobody had read all of them. I opened the epic planning to work through it top to bottom and closed it inside a minute. At that density there was no "working through it." There was only building something to read it for me.

First move was to actually count what I was looking at instead of guessing: 382 human comments, 139K characters, roughly 35K tokens total, median comment 206 characters. 180 of 382, 47%, had a pasted image attached. Only 67, 17%, carried an admin URL that named the page it was about. That last number mattered more than it looked. It meant text alone could auto-locate one in six items and nothing else.

At 35K tokens the whole corpus fit in a single context window. That changed the shape of the job: instead of an agent browsing around the product guessing at categories, it could read the entire pile once and propose a classification for every item, and I would confirm or correct rather than author each one from scratch. Confirming runs about 10x faster than authoring. That ratio is the entire throughput argument for doing it this way at all.

## The taxonomy has to come from the product, not from a planning doc

The instinct on a job like this is to sit down and invent categories: "bracket stuff," "pricing stuff," "misc." I tried a version of that first, sketching a taxonomy off title prefixes and label text pulled straight from the ticket corpus. It produced 40 distinct prefixes on open tickets alone, four spellings of "Tournaments," a bucket literally named `...`. Chunking on that would have meant every classified item landed in a bucket with no real destination, because the bucket wasn't a place, it was a guess about a place.

So I threw that out and drove the actual admin UI instead: staff login, `login-as` into a test account, Chrome over CDP via Playwright (chrome-devtools MCP kept dying mid-session, so the fallback was `connect_over_cdp` directly), walking every nav item and screenshotting it. The Tournament Manager area turned out to be a JS SPA where most sub-tab clicks never touch the address bar, it just stays on a single `?p=<module>&rand=...` URL while a JS call sets the active view and subview underneath. For those I called that state-setting function directly and recorded the `(view, subView)` pair as the node's `page_params` instead of a URL. First pass landed 34 nodes this way, 17 of them inside Bracket Manager alone, which is most of the product.

The payoff showed up immediately: comments carrying a `?p=` param that matched a node's `page_params` auto-routed for zero classifier cost, and that covered 17% of the corpus for free. Every node had a real path back into the product because it came from the product's own screens.

## The taxonomy still had holes, and screenshots lied about them

Five classification passes over 509 items still left roughly 49% with no fitting node. Not evenly distributed, either. It clustered in exactly three places: screens behind a button the crawl never clicked, screens gated behind tier or account type, and the entire visitor-facing site, which lives outside `/admin` so a nav crawl can't see it structurally.

That's when the second failure mode showed up. One batch's classification rules said flatly: "old screenshots can show UI that no longer exists." A 2023 comment had a screenshot of a "TEAM POOL DETAILS" dialog that had no live equivalent anywhere in the current product. The classifier had read the image correctly, matched it to something, and been completely wrong, because correctly reading an image is not evidence the screen it shows still exists. Worse example: a source-code check on the legacy ASP page behind the dashboard found the entire Dashboard screen (upcoming schedule, overview counts, bracket status panels) rendered as one static image tag over commented-out real markup. It was never built. Every item describing that "dashboard" was describing a design mock, and a classifier scoring by visual similarity alone would have routed all of them to a screen that had never once run in production.

That forced a rule change: image-based classification needs an explicit "this screen may be gone" check as its own step, separate from "does this image match a node." A follow-up pass added a re-capture (34 nodes grew to 84 plus 6 non-UI buckets, including a dedicated `never-built-prototype` bucket) and re-ran the parked items against it. Items that had genuinely gained a home this time got flagged that way explicitly; items that still matched nothing kept a `missing_node` field naming the gap instead of being forced into the nearest wrong bucket at high confidence.

## What came out the other side

Every item that couldn't be judged from text alone got its own disposition, REPRO, instead of being bulldozed into DO or DEAD by guesswork. And across all 82 tickets, a new `#Triage` field got backfilled so "how far along is this" stopped being a question you asked a person and became a query anyone could run against JIRA directly.

None of this made the backlog smaller by editing it down. It made the backlog legible by giving every comment an address in the actual product, derived from the actual product, not from whatever taxonomy sounded reasonable at 9pm.
