The feature had passed everything I could throw at it. The first real customer to use it broke it anyway.
What we were replacing
We had a prototype where any password worked, and the login screen said so outright. The job was to turn that into something real: an actual password, an actual rejection of the wrong one, and requests that would be remembered instead of vanishing the moment the page refreshed.
The design itself was ordinary and sound. A password gets scrambled by a well-understood method before it’s ever stored, with a setting that controls how many times that scrambling repeats. More repeats means more security, at the cost of a little more time to check a password. The number chosen for that setting was 150,000, comfortably above the 100,000 figure remembered from older guidance. Nobody checked whether the specific platform this would run on actually allowed a number that high.
Why the checks that ran before launch missed it
The setting only matters the moment a real account gets created on the real system. A build check does not create an account. Neither does a type check. Both ran clean, because neither one ever reaches the one line that actually depends on the platform’s real limit.
The first real signup did reach it, and it failed. The platform’s actual limit for that setting was 100,000, not 150,000. It did not quietly cap the number and move on. It refused the request outright.
The fix, and the test that actually mattered
The fix itself was a single number, 150,000 corrected down to 100,000, with a comment next to it explaining why it can’t go any higher. Then came the part that mattered more than the fix: a full run through the real system, not a simulated one. Create a new account. Try the wrong password. Try the right one. Save a request and check that it’s still there after a reload. All four of those checks ran against the live thing, not a stand-in for it, and all four passed.
What I’d tell myself beforehand
A platform’s limit is a fact about that platform, not something to carry over from a different job and assume still applies. Checking it takes a minute. The bigger habit is this: a test that never reaches the step that could actually fail has not proven anything about that step, no matter how many other things it did prove.
The number that matters now is 100,000, with a comment next to it explaining why, and the test that matters now is the one that runs against the real account, the real password, and the real reload.
AI Skills
Use this lesson with the AI assistant you already use
A login feature passed every check that could run without touching the live system. The first real signup hit a platform limit nobody had checked, because a build and a type check cannot execute the one line that actually depends on the real thing.
Paste the prompt, share only the context needed to answer it, and treat the result as a draft for your review. Do not include confidential information or let an AI assistant make changes without your approval.
Optional: for a visual report and saved memory, run /dxdev first.
Don’t have it? Get it at dxdev.com/skills/dxdev. The prompt works without it.
dxdev LESSON · paste into your AI coding agent
LESSON: A clean test is not proof until it exercises the one real step that can fail
SOURCE: dxdev.com/blog/2026-07-09_every-check-passed-the-first-customer-broke-it
WHAT HAPPENED: A password login feature was built with a setting chosen from general memory rather than checked against the actual platform it would run on. Every check available before going live passed, because none of them executed the one real step that depended on the platform's actual limit. The first real signup hit that limit and failed. The fix was one number, corrected against the platform's documented limit, followed by a full re-test that walked through account creation, a wrong password, a correct password, and a saved request, against the real, live system.
THE RULE: A setting chosen from memory or general guidance is a guess until it is checked against the actual system it will run on. A test that never executes the one step that depends on that setting has not tested it, no matter how many other checks passed.
CHECK MY CODE, then report PASS or FAIL with file:line for each:
1. Any setting or parameter chosen from memory, habit, or general guidance without checking the specific limit of the platform or system it will actually run on.
2. Any feature marked "tested" where the test suite never actually executes the one step most likely to depend on a live, real-world limit.
3. Any signoff or release decision based on a build passing or a type check passing, when neither one runs the code path in question.
THEN PRINT: a table (check, PASS/FAIL, evidence, fix) + a verdict (applies / partially / OUT_OF_SCOPE / no) + the single most important next action.