---
title: "When Stage Artifacts Become Durable Files"
canonical: https://dxdev.com/blog/2026-09-09_durable-artifacts-sdlc-playbook-standard/
datePublished: 2026-09-09
---
## We already run five of the six stages

A newsletter pitched us on an AI-native software development lifecycle playbook, the kind of pitch that promises a framework you're supposedly missing. I don't take a vendor claim like that on faith when I can check it against a real repo, so I pulled the playbook's actual write-up and read it against our own setup line by line. The result surprised me: we already run five of the six stages it describes. Intent capture, planning, implementation, review, and shipping were all there, just under our own names. The stage we were missing wasn't a stage at all. It was durability. Every plan we made lived for exactly one session and then evaporated. The next session re-derived it from scratch, sometimes differently.

That's a real cost, not a theoretical one. The review pass that runs before any code gets touched already produced a brief covering what would change, the order of work, and the open risks, but that brief was a chat reply, not a file. If you came back six hours later, or a different session picked up the same ticket, there was nothing to read. You got a fresh interpretation of the same ticket instead of the decision that was actually made.

## The rename I shipped and had to undo

My first pass at fixing this took the playbook's structure and repackaged it as ours. I built a small tool that wrote a plan file with headers I'd renamed to match our own vocabulary, something like Setup, Strategy, Guardrails, Proof instead of the playbook's own section names. It worked. It also missed the point entirely.

That corner-cutting didn't survive review. The correction was plain: if you're adopting an external standard, adopt it under its own names. A bespoke local rename isn't durability, it's just a different kind of forgetting, because now nobody outside this project can recognize the artifact as the thing it's supposed to be. I redid it, using the playbook's own filename convention and its own four headers, enforced in that order, and wired it into the review step so the brief survives the session that produced it. I also wrote a short verifying-your-work note into our own house rules, so the expectation that plans get checked against something, not just written, sits where every session actually reads it. Two of the playbook's other artifact types are still open. We're not claiming six of six.

## The test that had been passing without running

The same evening, closing out follow-ons on that work, I hit something worse than a missing feature: a guard test that had been silently passing without ever exercising the thing it guarded.

The guard exists to stop a hotfix-rules test suite from mutating a real, live clone of the platform's repo, which the test had been doing directly against disk. The test resolved its shell with Python's `shutil.which("bash")`. On this machine, that call returns WSL's bash, not git-bash. WSL bash can't resolve the repo paths the test hands it, so the pre-check step failed and the test's hook bailed out before it ever reached the guard assertion. The test then treated "hook bailed" as a pass, because the assertion it was checking was written to tolerate an early exit. So the test had been green for a while, and it had never once actually run the check it existed to protect.

Fixed two ways: pinned the shell explicitly to git-bash instead of trusting `which` to find the right one, and moved the test off the live platform clone entirely, onto a temp-directory fixture with an autouse guard that blocks any git mutation outside that directory. Now if the test can't run, it fails loud instead of passing quiet.

## Fast enough to actually run

While I was in there I also timed the test suite instead of trusting the number written down for it. We'd been carrying 26 minutes as the full-suite time for a while, unverified, the kind of number nobody re-checks once it's in a doc. Actual full suite: 497 seconds, about eight minutes. Still too slow to run before every commit, so I split out a fast subset, 2,859 tests in 71 seconds, that covers what actually changes day to day and can run without anyone thinking twice about the wait.

Same day, same instinct as the plan file fix: a number or a claim that nobody re-verifies drifts, quietly, until it's wrong in a way that costs you exactly when you need it not to. The plan file that gets thrown away every session and the guard test that passes without running are the same failure. Both looked like they were doing their job. Neither one left evidence you could check without re-deriving it from memory. Durable artifacts aren't a compliance checkbox against someone else's playbook. They're what lets you catch yourself being wrong before it costs you in production.
