984 of the 1,024 pages we imported from the old staff wiki had not been edited in five years, and the oldest was last touched in 2009. That number shaped the whole feature. It meant the hard part was deciding which of those pages anyone should believe.

What the old wiki turned out to be

The source was a MediaWiki 1.15 install that staff had been writing in for over a decade. Before writing an importer I counted what was in it. It was Semantic MediaWiki, so it wasn’t static text. 851 pages carried [[Property::Value]] annotations, and 358 pages rendered live {{#ask:}} tables from them, 562 queries in all.

That changed the plan. The importer captures each page’s properties as frontmatter and leaves a placeholder where a query block stood. Rebuilding the #ask behavior became its own ticket, a small resolver that reads frontmatter and answers queries, modeled on Obsidian Dataview. The import ticket stayed an import.

The importer is an owner-run one-shot script in the Node API. It writes straight to the shared database through the same query module the API uses, so it needs no minted token. Pages convert from wiki HTML to markdown with turndown. Media embeds get rewritten to /api/v1/kb/assets/:id. Writes are idempotent: pages MERGE on slug, and assets dedupe on sha256. Everything lands quarantined with status='imported'.

The sandbox run that had to be torn down

The first pass ran in sandbox mode, with KB_SANDBOX=1 stamping sandbox=1 on every row so reads filtered them out. That was the right instinct and it still cost us a step. The slug column is globally unique, so sandbox rows and live rows cannot coexist. Going live meant deleting all 1,024 sandbox rows first.

The live command needs three flags: --write --live --i-mean-it. The result was 1,024 of 1,024 pages, 0 failed, 319 assets. 27 of those assets are not images (17 PDFs, 7 mp3s, 3 xml files), so they get linked instead of embedded as <img>. A rollback SQL file was written before the run, and we never needed it.

Gating the import on a review ladder

Importing 1,024 pages into a knowledge base and calling them knowledge is how you get a wiki nobody trusts. So the import was gated on a review ladder shipping first. Every imported page starts at reviewTier=unreviewed. A staff member moves it along: in_review, then verified, with a staleness date attached. Pages can also be marked needs_update or archived.

The API exposes it two ways:

  • POST /pages/:slug/review performs a transition and is staff-only.
  • GET /review/queues returns counts per tier. GET /review/queue/:name returns one queue, oldest source edit first. The names are unreviewed, in_review, needs_update, archived, and due.

due is the interesting one. A nightly job expires verifications that have lapsed, and those pages fall back into a queue instead of staying green forever. Oldest-first ordering is why the 2009 pages sit at the top and get judged first.

The importer seeds review data as frontmatter, and the ladder wants real columns. So after the live run, a backfill script parses the frontmatter and stamps the review columns, quoted timestamps included. I verified in the production database: 1,024 live rows, all imported and unreviewed, sourceLastEdited filled on every one, sandbox count zero.

The agent that said there was no UI

With the import verified, I asked an agent to close out the ticket and tell me where I could look at the result. The ticket had no link. I wanted one, because that is the point of having a developer space.

The agent searched the API repo. It found the KB routes, saw that every handler ended in res.json, found no view templates, and reported that no viewing UI existed. It proposed reopening the ticket with an honest status and spinning the missing page-view UI into a new ticket.

I nearly signed off on that. I didn’t have the right memory to challenge it, only a vague one: I remembered fixing view issues and syncing views on an earlier ticket, so I pushed back and asked how there could be no UI. The agent’s first response was to re-verify, and it said plainly it was doing that because it did not want to repeat a wrong claim.

The UI existed. It lives in the staff SPA repo, not the API repo: KbListPage, KbEditorPage, KbReviewQueuesPage, and a ReviewBadge component with the transition buttons, all merged to develop and wired into the router. The agent had searched one repo out of several and treated the absence of evidence there as evidence of absence.

The cost was real. I burned about ten minutes of a working afternoon being told the feature was unbuilt, and I was one reply away from creating a duplicate ticket for work that had shipped. The ticket also went out with a status I had to unwind. The check that would have caught it is boring: before you assert “X does not exist”, say which repos you searched. Ours span an API, a staff SPA, and several legacy clones.

Auditing the import instead of trusting the ticket

Since the ticket had no link to the browser surface, I stopped trusting it and ran an adversarial audit of the import. That job was read-only against the live database. It sampled across all 1,024 pages and was told to quantify defect classes: dead internal links, media fidelity, the frozen Semantic MediaWiki content, table conversion, slug collisions, junk pages. For side-by-sides it logs into the live wiki and renders the same page next to the imported markdown.

I also rendered a batch of imported pages straight from the database to HTML, because that is the fastest way to find out whether “0 failed” means “correct”. A clean exit code from an importer means the rows landed. It says nothing about whether a page reads correctly.

That is what the ladder is for. Nobody, human or agent, has to certify 1,024 pages at once. Each page earns its badge, the oldest and likeliest to be wrong get looked at first, and a verification expires on schedule instead of quietly going stale again.