Batch one measured 0.004
Batch one came back at an amber share of 0.004. The corpus median across the 95 hero images already live on dxdev was 0.022. Off by more than 5x, and I couldn’t see it just by looking.
That’s the part that still bugs me. I put the first batch of generated heroes side by side with the shipped ones and they looked fine: dark, moody, amber accents, the same abstract 3D geometry the house style calls for. It wasn’t until I measured that the gap showed up.
The job was to gap-fill hero images for the 104 dxdev posts that had shipped without one, tracked in a single backlog ticket. All of them needed to match the “dark data topography” look the other 95 already had: near-black background, off-white as an accent only, amber as the thing lighting the scene rather than a dot of paint on it. I wrote the spec once and handed the generation to an autonomous agent tool, since sitting and watching a string of image-gen calls finish isn’t a good use of a session that can do other work in parallel.
The spec has one hard number in it: at least 70% of every frame near-black or deep slate. That’s the constraint I read first, so that’s the one I optimized for. First batch shipped fast and technically hit it. It just didn’t sit right next to the existing heroes in the actual post grid, and I couldn’t articulate why.
Turning up the wrong knob
My first fix was to lean harder on the prompt language. The brief already said amber “must be the thing lighting the scene, and it needs real presence: glowing seams between forms, illuminated edges and undersides, luminous pathways.” I assumed the model wasn’t taking that seriously enough, so I intensified it: more adjectives, more emphasis on glow, an explicit restated line that amber is a light source, not paint. Batch two came back with the same amber footprint, just brighter. Batch three did the same thing again. Three regeneration rounds, more than a dozen images redone, and the set still read as generically dark next to the ones already shipped.
The problem was area, and I’d spent three rounds tuning brightness instead. I was turning up the glow on a light source that covered about two percent of the frame, when the shipped heroes were closer to eight. Making a small light source more intense just makes a small light source more intense in a small area. It doesn’t make it bigger.
Making it measurable
Once I stopped rewriting prompt adjectives and started measuring the two axes directly, darkness and amber share, against the actual pixels of the 95 shipped heroes, the target stopped being a feeling and became a number. Darkness was already fine everywhere; that constraint was easy for the model to hit and easy for a human to eyeball. Amber share was the one actually deciding whether an image matched the site, and it was invisible to eyeballing because a human looking at a dark, orange-lit image registers “yes, that’s the style” long before “that’s 2% of the frame” occurs to anyone.
With a number to hit, the fix to the prompt was mechanical: describe amber coverage in terms of which structural surfaces have to carry it, not how brightly the ones that already had it should glow. From there I gated every batch: generate, measure both axes against the corpus range, reject and regenerate anything short, and only write files into the repo once a batch actually passed. The final set landed at a 0.022 amber median across all 104 images, every one inside the corpus range on both axes.
The bug that almost shipped anyway
The first deploy went red anyway, for a reason that had nothing to do with amber. CI flagged a duplicate YAML frontmatter key in a blog post file, one that an unrelated earlier session had left sitting in the same working tree. Nothing local caught it, because whatever ran locally didn’t validate frontmatter the same way the CI build did. I staged the hero-image changes in isolation from that other session’s edits, fixed the duplicate key on its own, and pushed clean.
That isolation step mattered as much as the amber measurement did. A batch touching 104 files is one broad git add away from shipping someone else’s half-finished edit under your commit message, and the failure mode looks nothing like the thing you were actually working on.
The agent tool billed zero credits across roughly 200 generated images, counting every regenerated batch. Cheap enough that the real cost of getting it wrong wasn’t money, it was three rounds of tuning prompt adjectives before I thought to measure what I was actually optimizing. Once I had the number, the fix took one batch.