---
title: "Silent Integration Failures: When Your Webhooks Disappear Without Warning"
canonical: https://dxdev.com/blog/2026-08-22_silent-integration-failure-detection/
datePublished: 2026-06-22
---
At 2:07 PM on June 22, we fixed a payment that had been charged but had never reached the account that paid it. The customer had done their part. Our recurring-payment integration had not.

The account had expired because the payment record never appeared in our application. The charge was real. The renewal was not. We corrected the account, applied the missing payment, extended the term through August 15, and added a free month for the trouble. Then we opened the incident instead of treating it as a one-off support mistake.

That decision uncovered a worse problem. Since late May, an edge proxy had been blocking posts from the recurring-payment gateway before they reached our webhook handler. Customers could pay, the gateway could accept the charge, and our system could receive nothing. There was no application exception because the application never got a request. There was no alert because we had never defined a missing stream of payment events as an alertable failure.

The integration did not fail loudly. It failed by going quiet.

## The path from an expired account to the edge

The first symptom looked like a normal reconciliation issue. A customer had a charge in June, but their account still showed expired. We had three plausible places to look.

The first was the account update itself. A valid payment might have arrived and a renewal job might have failed afterward. The second was our payment-recording path. The handler might have accepted the webhook but failed before committing the payment. The third was delivery. The gateway might have charged the customer but never completed the request into our system.

Those are distinct failures, and they leave different evidence behind. A renewal bug leaves a recorded payment with an unchanged expiry date. A handler failure leaves a request or application error. This incident had neither. It had a real charge and no corresponding internal payment event.

The missing payment ruled out the cleanest version of an account-state bug. The absence of a handler-side record pushed the investigation outward. We traced the gateway's posts through the public route and found that the edge proxy had been silently blocking them. The request was stopped before it crossed into the application boundary.

That distinction matters. We had monitoring around code that ran after a request reached us. We did not have health monitoring around the fact that a critical outside system was still reaching us at all.

A blocked webhook is not necessarily an application error. It can look like nothing.

## Why the existing signals were useless

We had support tickets, account inspection, and payment reconciliation. They are useful after someone has found a bad outcome. They are not integration health checks.

The delivery path crossed the gateway, the internet, our edge proxy, the public route, the webhook handler, the payment record, and the account update. The block sat outside the part of the chain that produced normal application logs. Handler-error monitoring could reveal a malformed payload or failed database write. It could not reveal that the handler had received zero valid callbacks for weeks.

The customer ticket could not be the detector. An expired account after a valid payment is a customer-impact signal, not an operational control. By then, we are repairing trust as well as data.

The incident took 2 hours and 15 minutes to resolve from the initial payment report. The root-cause investigation itself took another 1 hour and 17 minutes. That time was not spent on a difficult code fix. It was spent proving which boundary had gone silent and establishing how long it had been silent.

## The monitor needs two sides

The monitoring follow-up we filed is not a generic "webhook failed" alarm. A single request failure is noisy. Gateways retry. Networks flap. Some payment schedules naturally produce uneven traffic. An alarm that pages on every non-200 response becomes a notification people learn to ignore.

The pattern is a two-sided health check. One side measures the outside fact that matters, payments settled by the gateway. The other measures the inside fact, recurring-payment events recorded by our system. The monitor compares the two over the same window and investigates a gap.

```yaml
integration: recurring-payment-webhook
observe:
  outside: gateway recurring payments accepted
  inside: payment events recorded by the application
alert_when:
  expected_events_are_missing: true
  outside_and_inside_counts_diverge: true
context_to_capture:
  last_successful_delivery
  public_route_response
  edge_proxy_decision
  affected_payment_window
```

That configuration is deliberately about evidence, not a magic number. The normal volume and payment cadence determine the threshold. A system with continuous recurring activity can alert after a short quiet period. A system with sparse schedules needs a longer window and a reconciliation check. In both cases, the rule is the same: alert when a payment-producing system and a payment-recording system stop agreeing.

The second signal is important because it catches more than edge blocking. It would also surface a disabled route, an authentication change, a certificate problem, a handler that accepts requests but does not persist them, or a gateway configuration drift. The failure modes differ, but the customer-visible outcome is the same: money moved and the account did not.

## The alternatives I rejected

I considered treating this as a route-specific fix. Restore the blocked delivery, write down the edge-proxy behavior, and move on. That would repair the exact incident but preserve the detection gap. Another edge rule, an upstream configuration change, or a future webhook integration could recreate the same class of failure.

I also rejected a monitor that only watches application errors. It gives false confidence at precisely the boundary where this incident happened. Zero requests can look healthier than a burst of errors if the dashboard only displays failures from code that ran.

A scheduled manual reconciliation was the third option. It would have detected mismatches eventually, but it would still make the team discover a live break by batch review. The missing-event monitor is better because it detects the integration losing its pulse while there is still time to restore delivery and make affected accounts whole quickly.

We restored delivery, corrected the affected customers, documented the edge-proxy gotcha, and created the monitoring work so recurring payments cannot quietly disappear again. The fix was not simply allowing one POST through one boundary. The fix was deciding that silence, in a system that should be receiving money events, is a failure state with its own evidence and its own alarm.

A webhook route is not healthy because it has no errors. It is healthy because the events that should arrive still arrive, get recorded, and produce the state change customers paid for.
