---
title: "When Enforcement Scatters, Data Silently Leaks"
canonical: https://dxdev.com/blog/2026-09-18_weak-link-validation-cascade/
datePublished: 2026-09-18
---
Twenty-two customer tickets closed without a sprint, and only one of them could still be repaired. The other 21 lost that field for good.

The rule was simple: every customer ticket goes in the current sprint. Our support flow runs through a Discord bot that creates and closes tickets in JIRA, plus some agent tooling that does the same. Sprint is how we count what got handled in a given period, so a ticket with no sprint vanishes from the numbers. I noticed it when the counts didn't add up, and the pattern was plain once I looked: closed tickets, empty Sprint field, and the bot behind most of them.

## Four guards on the wrong thing

The obvious suspect was the bot, so I started there. It created tickets and moved them through their transitions, and it was the biggest source of the bad closes. Fixing it was the right first move, and it was also where I went wrong, because I treated the bot as the whole problem.

I made the bot refuse to close a ticket without a sprint. Then I did the same in the agent tooling, since the agents hit the same transitions. Then I wrote a JIRA validator, and finally a daily watchdog that sweeps for closed customer tickets with an empty Sprint and flags them. Four layers, each verified against a real ticket, each one green.

That felt thorough. What it actually was, was four copies of one rule living in four places, each guarding a caller. None of them guarded the thing the callers were calling.

## The unlocked door in the workflow

Once the four fixes were in, I went back to the JIRA workflow itself and read the close transition, the one we use for tickets closed with no code change. Its screen listed Sprint as optional. That was the whole bug. Anything that could reach that transition could close a ticket with no sprint, and nothing in JIRA would object. A person in the web UI, a script, a future tool I hadn't written yet. The bot had just been the most frequent visitor.

The four layers were guards around the door. The door itself was still unlocked, and it had been unlocked the whole time the tickets were leaking.

The cost of that ordering is concrete. Every ticket that closed through that transition before the watchdog existed had no safety net behind it. The watchdog only sees what's still recoverable when it runs. By the time I was auditing, 21 of the 22 had lost their sprint data permanently. Only one was still repairable. I filed three follow-up tickets for the leftovers so each gets a clean scope and a clean history, and wrote a paste-ready handoff so the next session doesn't have to reconstruct any of this.

## The order I should have worked in

If I were fixing it fresh, the order flips. Make Sprint required on the transition first, in JIRA, where it applies to everyone. Then decide which of the other layers still earn their keep.

- The transition validation is the enforcement. It runs regardless of who calls it.
- The bot check is still worth having, but only for a better error message. A refused close from JIRA is cryptic in Discord, and a check in the bot can say "add a sprint" in plain words.
- The watchdog is a tripwire. It catches anything that slips past the transition, and it tells me the enforcement has a hole. It should almost never fire.
- The agent tooling check is the same story as the bot: a friendlier failure, not a second source of truth.

Written that way, the other three are convenience and detection, and the one rule lives in one place. When I have to change it, I change it once.

## Why the fifth path stayed hidden

Optional-by-default fields are quiet. Nothing errors, nothing logs, and the data just isn't there. You find out weeks later when a count is off, and by then the information you need to backfill it may not exist anywhere. Scattered enforcement makes this worse in a specific way: every new caller you add is another path that has to remember the rule, and the path you forget is the one the data leaks through. I had four paths covered and was confident about it. The fifth path was the workflow, and I only found it by reading the transition config instead of reading my own code.

Before I add a check to a caller, I now ask where the state change actually happens, and whether the rule can live there. If it can, that's the fix. Everything else is a convenience on top of it.
