Someone sent me a screenshot with one line under it. A task kept skipping the current work period. The person it should have been assigned to was blank. And if there was a version number involved, that was blank too. Not once in a while. Every single time.

I had a tool built specifically to fill in those three details automatically the moment a task got marked done. It had been sitting in the system for a while. Nobody had complained about it, so I assumed it was doing its job quietly in the background, the way a good automatic step should.

It had never once worked. Not on that task, not on any task, not for anyone, since the day it was turned on.

The code was fine. The wiring in front of it wasn’t

Here’s the part that got me. The logic inside that tool was genuinely well thought out. It knew how to tell the difference between everywhere a task had ever passed through and where it currently sat. It handled the messy cases correctly. Somebody had clearly sat down and solved the actual problem.

And it had never run. Not because the logic was wrong, but because something earlier in the chain was broken in a way that stopped the whole thing before it ever got to the good part. One small setup error at the very front of the tool crashed it before a single line of the careful, correct logic underneath ever had a chance to fire.

Once that first crash was fixed, it turned out there were four more places, stacked one after another, where a single missing piece could have silently killed the whole thing again. Any one of the five would have produced the exact same result: nothing happens, and nothing tells you that nothing happened.

Why “nobody complained” isn’t proof

This is the trap. A tool gets built, it gets tested once, the test passes, and it ships. From that point on, the only signal anyone has that it’s still working is silence. No error messages, no complaints, nothing flagged.

But silence isn’t a good signal. Silence is what you’d see whether the tool was working perfectly or had been broken since day one. The two situations look identical from the outside. The only way to tell them apart is to go check a real, recent case and see whether the thing that was supposed to happen actually happened.

I keep three details I actually care about on every closed task, and I found out all three had been broken the whole time, from the same screenshot, on the same day, because I happened to be looking at that particular task closely enough to notice.

Where AI fits

This is exactly the kind of checking an AI assistant is good at, if you point it at the right question. Not “does this look like it’s set up correctly,” but “did this actually happen on a real, recent case.” An assistant can trace what a step was supposed to do, pull a handful of recent examples, and tell you plainly whether the expected result shows up or not.

What it shouldn’t do is turn anything on or off, or decide on its own that a fix is safe to apply. Its job is to surface the gap between “this was built to happen” and “this actually happened,” and hand that back to a person to act on.

The human decision

A person has to decide what “proof” looks like for each automated step that matters, and how often it’s worth checking. Some things are worth a weekly glance. Some things, especially anything touching money, customer records, or a deadline, deserve a standing check that flags itself the moment it goes quiet, rather than waiting for someone to happen to notice.

The lesson

A tool passing its first test and getting turned on proves it worked once, in that one moment. It does not prove it has worked since. If something is supposed to run automatically and matters to you, don’t just trust that it’s still doing its job. Go find a recent, real example and check.

The paired Build Log walks through how five separate, stacked failures let a well designed piece of automation sit in the system doing nothing, and why the fix that mattered most wasn’t the code, it was building a way to know the difference between “this exists” and “this ran.”