---
title: "Turning a Stack of AI Transcripts Into Beats for Under a Dollar"
canonical: https://dxdev.com/blog/2026-04-30_prompt-caching-turned-transcripts-cheap/
datePublished: 2026-04-30
---
Five conversation transcripts added up to 858KB of Markdown, roughly 215,000 input tokens, all headed into the same ingestion pass, each one about to repeat the exact same system prompt in full.

That repetition changed the job.

I was building the first ingester for a content system that turns raw AI conversations into publishable candidate beats. Wave 1 is deliberately narrow. Each transcript becomes one candidate arc. The agent reads it, decides whether there is a publishable beat, proposes tags and a slug, then writes structured records to Postgres. Multi event clustering can wait for Wave 2.

The input was not a polished brief. It was five full conversation exports. They include the failed starts, tool output, revisions, and the parts that only make sense when you can see what came before. We did not want an operator reading every file before the model could begin, and we did not want a vague summarizer producing prose that was difficult to route, filter, or reject.

## The bill was in the prompt, not the transcript

The model gets the same system instructions for every transcript. They define the editorial test, the required fields, rejection behavior, and the constraints on the eventual output. The transcript itself arrives as the user turn.

Five conversations meant repeating the full system prefix five times if nothing were done about it, even though nothing in that prefix changes between calls. The source material is the expensive part by volume, but the repeated instructions would still have been needless spend on top of it.

So the ingester marks the system prompt as cacheable from the first request. The first call establishes the cached prefix. Requests two through five reuse it while the transcript remains the variable input. It is not magic. The user turn still contains 858KB of raw Markdown, so transcript bytes dominate the token count by volume. It does mean the fixed system-prompt tokens are paid for once instead of five times.

Sonnet 4.6's published rate card is $3 per million input tokens and $15 per million output tokens. Run 215,000 input tokens through that rate and it works out to about $0.65, plus a small output charge on top, for a full pass of roughly $0.73. Round it honestly, and the whole batch is a sub-dollar experiment, not a recurring line item.

## A schema beat the convenient alternatives

The first discarded option was plain text summaries. They would have been cheap to wire up, but every downstream consumer would need to interpret prose again. A later step would still have to decide whether a transcript was publishable, extract tags, assign a unique slug, and preserve the provenance of the run. That moves uncertainty through the system instead of containing it at ingestion.

The second option was a permissive JSON response. That looks structured until the model omits one field on a rejection, changes a field type, or returns a half formed object after a long transcript. Then the database write path becomes a repair shop.

I used strict schema-enforced output instead. Every field is required. When a transcript has no publishable beat, the model returns empty strings and arrays for the fields that do not apply. That feels slightly awkward in the response, but it makes rejection mode explicit and keeps the writer simple. One skeleton beat goes into the beats table, one record of the run goes into the ingestion jobs table, and any suggested tags create the corresponding tag records.

The other important field is not generated at all. The transcript's own identifier is the idempotency check. On a rerun, the ingester looks for that transcript ID first and skips work already recorded. Slugs follow a conversation-start-date-plus-proposed-slug pattern, adding a numeric suffix only if there is an actual collision.

## I debugged the wiring without buying model calls

Before sending 215,000 tokens anywhere, I ran the ingester in dry-run mode. It parsed the manifest, found all five paired transcripts, and printed the work it expected to do. No model request and no database write. A limit flag gives the same control when I want to test a smaller slice.

That dry run caught the distinction that mattered. The pipeline wiring was ready. The live pass was not failing on parsing, schema validation, cache configuration, or the database path. It was blocked because the API key for the model provider was not present in the environment. That is a setup task, not an ingestion bug. I ran the actual sweep by hand instead, one transcript at a time, and fed the results through the same write path the automated ingester will use once the key is in place.

Once the key is available, a changed system prompt is no longer a reason to protect an old batch result. I can rerun the five transcripts, compare the new skeleton beats and tags, and decide whether the editorial instructions improved the extraction. At that per-token rate, iteration is cheaper than pretending the first prompt was finished.
