Three hours and 40 minutes into a tournament Teams-tab prototype, I had a working branch that deserved to be deleted.
The prototype was not broken. It had a bracket versus list-view toggle. In bracket mode, an Available Teams rail sat beside the real bracket diagram, and the advancement slots were grayed out. I could run it on a local clone and show the behavior end to end. The problem was worse than a bug. I had built a coherent answer to a question nobody had asked.
I started from inference instead of opening the designer mockups.
The direction I invented
The ticket was a feature overhaul, not a small defect. That should have changed the first step. Instead, I read the ticket language, looked at the existing tournament flow, and let the agent work forward from the codebase. The result was a plausible interaction model: give staff a way to move between a list and a visual bracket, keep unplaced teams available in a side rail, and make unavailable advancement positions visibly inert.
That model was internally consistent. It was also my model.
There were two ways to begin. The first was to pull the designer screenshots into the task context and implement the shapes, hierarchy, and interactions they specified. The second was to infer the intended screen from the existing application, the task wording, and the data already visible in the bracket. I chose the second route because it felt faster. The code was nearby. The expected feature sounded familiar. The agent could start immediately.
That was the wrong optimization.
Existing code can tell an agent what the system currently permits. A ticket can tell it what must change. Neither source can reliably tell it how a redesign is supposed to resolve the choices that remain open: which control carries the primary action, which state deserves visual weight, what moves together, or what a user should notice first.
A mockup is not decoration added after the architecture. For an interface overhaul, it is part of the specification.
The local demo that failed its real test
The diagnostic path was uncomfortable because the branch behaved exactly as I had built it to behave. I did not find a failing command, a bad response, or a missing record. I opened the local clone, exercised the toggle, inspected the bracket view, and saw the available-teams rail beside the diagram. The inactive advancement slots were gray. The view could be demonstrated.
That made the failure easy to misclassify. I could have called it a polish gap and spent another session adjusting spacing, labels, or edge states. I could have treated the visual feedback as a request for incremental revisions. Both choices would have preserved the original mistake: the implementation had been derived from inference, not from the design artifact that defined the work.
We stopped before that happened. On review, we decided to abandon the direction rather than negotiate it into compliance. The local branch was deleted. The ticket was reset to preparing. The next session would begin clean, from the mockups.
Deleting a branch is often described as waste. In this case it was containment. The longer I kept the prototype alive, the more likely its assumptions were to turn into constraints. A functioning UI has persuasive force. Once people can click it, it becomes tempting to ask how to save it. That is a dangerous question when the visual model is wrong at the root.
What I changed in the agent handoff
The corrective action was not a better prompt that said, “be more faithful to design.” That leaves the agent with the same incomplete evidence and asks it to guess more carefully.
The handoff has to name the source order. For a visual overhaul, the agent needs the mockups first, then the ticket’s acceptance criteria, then the existing code. The implementation plan should trace each meaningful interface decision back to one of those sources. If the mockup does not answer a behavior question, the agent should surface the gap instead of silently borrowing a pattern from the old screen.
That sequence also makes review sharper. We are no longer judging whether a prototype seems reasonable in isolation. We can ask a much smaller and more useful question: where did this element come from? If the answer is a screenshot, a stated requirement, or an explicit design decision, we can review the implementation. If the answer is “it seemed like the natural way to do it,” the work is not ready to build.
I do not regret making the prototype. Seeing the bracket and the rail together made the error obvious in a way a planning note would not have. But I regret letting an agent reach for code before I gave it the visual source of truth.
The branch was coherent. That was never the standard. It needed to be faithful.