At 2:20 AM, I changed our release runbook because one release had exposed a drift we could not safely ignore.
The release itself was not the problem. The problem was that our process could look at the repository and produce two different answers to one question: what comes next?
For too long, the runbook derived the next release version from Git tags alone. That felt reasonable because tags are the visible release marker. A tag says that a commit became a release. It is easy to inspect, easy to sort, and easy to turn into the starting point for the next version.
It is also incomplete.
The merge trail held version information of its own, and that record had moved past what the tag check was reporting. The two histories had drifted. One of the affected releases was a teammate’s. The tag-only rule did not know it was looking at a lower answer than the release history could support.
That is a small failure on paper. In a release process, it is a bad one. Version derivation is supposed to be boring and deterministic. If the same repository can answer with a tag-derived value in one place and a merge-trail value in another, the tool is not determining the version. It is choosing which evidence to ignore.
The diagnostic was two answers, not one broken tag
The first useful observation was that this was not a single malformed release. It was a disagreement between the sources we were already relying on.
The old flow asked one question:
What is the highest release tag?The corrected flow asks two:
What is the highest release tag?What is the highest version represented in the merge trail?Then it chooses the higher version before it calculates the next release.
next_release = increment(max(highest_git_tag, highest_merge_trail_version))That is the whole change in one line, but it changes the failure mode completely. A stale or missing tag can no longer silently pull the next release backward if the merge trail has already recorded a higher version. The runbook now treats disagreement as input to reconcile, not as a reason to trust the first source it queried.
I did not replace the tag check with a merge-only check. That would have swapped one blind spot for another. Tags are still our explicit release markers. The merge trail is still evidence of release work that the tags alone might fail to reflect. Neither is authoritative enough by itself once we know they can disagree.
I also did not make the fix a manual cleanup rule. We could have repaired the known tag state around that release and told people to be more careful. That would have fixed the symptom in one repository state while preserving the same derivation bug for the next one. A release process needs a rule that survives the operator having an ordinary rushed day.
Local success is not release success
The second change was less glamorous and just as important. The runbook now verifies that the tag reached origin before it calls the release complete.
A locally created tag is not evidence that everyone else can see the release. It is only evidence that one working copy has it. If the tag has not reached origin, the repository has split release state again: the local machine believes the version exists, while the shared remote cannot confirm it.
So the revised sequence is deliberately ordered:
- Read the highest Git tag.
- Read the highest version in the merge trail.
- Take the higher of those two values and derive the next release version.
- Create and push the release tag.
- Verify the tag exists at
origin. - Only then treat the release as done.
The last verification is not a cosmetic receipt. It closes the loop on the evidence the next run will consume. The next operator should not have to infer whether a local action became shared repository state.
The runbook is now a guardrail, not a memory test
I updated the written release flow and sent the change, along with a sync script, to the teammate whose release had surfaced one of the drifts. I designed the original tag-only derivation. It was mine, and it produced a wrong answer for someone else’s release before it produced one for mine, which is a worse way to find out a rule is incomplete than catching it on your own work first. That review matters because release tooling is shared infrastructure. The person who sees the mismatch should be able to follow the same sequence, reproduce the decision, and check that the remote agrees.
This was a 48-minute fix, not a new release platform. We did not add another service or invent a database for version state. We made the existing evidence explicit, compared it before acting, and checked the remote after acting.
The important part is not that Git tags are unreliable. They are useful. The important part is that a release number is a conclusion drawn from repository history, and a conclusion is only as good as the records it considered.
Now our release flow has to look at both records. If they disagree, it takes the higher version. If it creates a tag, it checks origin. It was a 48-minute fix, one line of logic (next_release = increment(max(highest_git_tag, highest_merge_trail_version))) and one verification step, for a bug that had already shipped a wrong number once before anyone traced it back to the runbook.