---
title: "When Your GTM Trigger Breaks Your Funnel: Detecting Data Pollution"
canonical: https://dxdev.com/blog/2026-08-22_gtm-trigger-event-pollution/
datePublished: 2026-07-22
---
On July 22, I found an onboarding event that had fired on every page load since July 12. For 10 days, our signup funnel had been counting navigation as progress.

The event stream was not blank. It was busy. Dashboards rendered and reports produced percentages, and for ten days I read those numbers as a funnel doing fine. I did not stop to ask why the top of it looked unusually active, because unusually active looked like good news, not like a defect. Nobody flagged it because nothing about a healthy-looking chart asks to be questioned. We could have spent a week debating conversion behavior when the system was really measuring a tag manager mistake, and the only reason we didn't is that July 22 happened to be the day I looked closely enough to catch it, ten days after it started.

We had an event storm, not a sudden burst of healthy onboarding. A GTM trigger was attached broadly enough to fire the onboarding event on each page load. The signal we needed was a user reaching a specific onboarding step. The signal we had was the browser loading another page.

Those two things are not close substitutes.

## The funnel was not merely noisy

A broken event pipeline is often obvious when nothing arrives. This was harder because too much arrived. Every page load added another record that looked like a legitimate onboarding action. Once that happened, event volume stopped being evidence that people were moving through the flow.

The failure had two layers. First, the trigger manufactured events. Second, our funnel did not yet have the step dimension needed to distinguish real signup progress from a generic event count. The trigger was the source of the pollution, but the missing dimension made the pollution easy to mistake for behavior.

The useful question was, "What condition causes this event to fire, and does that condition represent the business action the event claims to record?" In this case, the answer was page load. A user could generate repeated onboarding events without advancing at all.

## Treat a rising event count as a hypothesis

When I am checking a funnel, I want the event name, its firing condition, and the dimension used for analysis to agree. If any one of those describes a different thing, the dashboard is a story generator.

The diagnostic path here started with the storm itself. We traced the onboarding event back to GTM and found the trigger firing on every page load. That explained both the shape of the stream and the 10-day gap between the July 12 break and the July 22 discovery. Nothing was technically down. The analytics path was accepting exactly what we sent it.

That distinction matters. Delivery checks answer, "Did the event arrive?" They do not answer, "Did the event mean what its name says?" A tag can successfully transmit a false fact thousands of times.

I now use a simple test: describe the event and its trigger in one sentence each. If those sentences are not semantically identical, inspect the event. "Onboarding step completed" cannot be triggered by "page loaded."

The record needs enough dimensions to make the test practical after the fact. For this funnel, we registered the step dimension so the signup path could finally be queried as steps rather than treated as one undifferentiated bucket of onboarding traffic. That did not repair the historical 10 days. It restored our ability to separate stages going forward.

## Containment came before cleanup

We shipped a deduplication hotfix first. That was containment, not a diagnosis. It stopped the junk stream from compounding while a teammate published the trigger correction. Within minutes of the trigger fix, the flood collapsed.

That sequence was deliberate. We could have relied on deduplication alone, but it would have left the misconfigured GTM trigger generating bad intent at the source. It also would have made the reporting layer responsible for compensating for a tag-management error. The correct architecture is the opposite: the trigger emits one event for one business action, and downstream deduplication is a guardrail for retries or accidental duplicates.

We also pushed a full legacy-tag cleanup into the release train. That work matters because stale tags create a second failure mode. Once several tags can describe roughly the same action, nobody can tell whether a spike comes from behavior, an old container rule, or two instruments reporting the same thing under different names. The cleanup reduced the number of moving parts we would have to inspect the next time a chart looked implausibly healthy.

## Recovery means declaring the bad interval

The 10 days of polluted data cannot support claims about onboarding behavior. We mark the interval as unusable for funnel interpretation. The stream remains useful forensic evidence because it tells us when the trigger began firing and why the funnel appeared active. It is not product signal.

After the fix, I validate three things in order. First, the GTM trigger only fires when the named action occurs. Second, one action produces one event. Third, the step dimension makes the funnel queryable as a sequence. Each check covers a different failure point.

The event stream quieted within minutes of the trigger fix landing on July 22. The 10 days from July 12 to July 22 stay marked unusable for funnel interpretation, permanently. That is the actual size of the mistake: not a chart that looked wrong, but ten full days of data nobody can ever use to answer what onboarding behavior actually looked like during that window.
