Engineering
Webhook Delivery Failures: Causes, Detection and Recovery
Webhooks fail quietly and asymmetrically, the sender believes delivery succeeded while the receiver never processed the event. A practical guide to the failure modes, how to detect them, and how to recover.
A webhook delivery failure is any case where the sender considers an event delivered but the receiver never successfully processed it. The asymmetry is the problem: the sender's dashboard shows a 2xx response, so nothing looks wrong from either side until downstream data is found to be missing.
- A 2xx response means received, not processed. Acknowledge only after durable persistence.
- The four dominant failure modes are silent 2xx, retry exhaustion, ordering violations and signature drift.
- Recovery requires a replay path. If you cannot re-request a window of events, you cannot recover from an outage.
Why webhooks fail asymmetrically
With a polled API, the consumer controls the request and sees every error. With a webhook, the producer controls delivery and sees only the HTTP status your endpoint returned. If your handler returns 200 and then throws while writing to the database, the producer's records say delivered and yours say nothing happened.
That asymmetry is why webhook failures are so often discovered by a customer rather than by an alert.
The four dominant failure modes
Silent 2xx
The most common and most damaging. Your endpoint acknowledges before the work is durable: it returns 200 and then enqueues, parses or writes. Anything that fails after the acknowledgement is invisible to the sender.
Fix: acknowledge only after the event is persisted somewhere you can replay from. Write the raw payload to durable storage first, return 200, then process asynchronously. The acknowledgement should mean "I will not lose this", not "I have finished with this".
Retry exhaustion
Providers retry on a backoff schedule and then stop. The windows vary widely: some give up after a few hours, others persist for a day or more. If your endpoint is down for longer than the provider's retry window, those events are gone unless there is a replay API.
Fix: record each provider's retry policy and maximum window in your integration inventory, and know before an incident whether a replay path exists.
Ordering violations
Webhooks are not ordered. A updated event can arrive before the
created event it depends on, particularly after a retry. Handlers written as if
order were guaranteed corrupt state in ways that surface much later.
Fix: make handlers idempotent and order-independent. Key on the resource identifier plus a version or timestamp, and discard events older than the state you hold.
Signature and secret drift
A rotated signing secret, a changed signature algorithm, or a payload serialisation change causes every delivery to fail verification at once. This one is at least loud, but it is often misdiagnosed as an outage on the provider's side.
Detecting what you cannot see
Because the sender's view is unreliable, detection has to be built on your own expectations:
- Expected-volume alarms. Alert on the absence of events. If a webhook normally delivers several hundred events on a weekday morning and delivers eleven, that is a detection even though nothing errored.
- Sequence gap detection. Where the provider supplies a monotonic identifier, track gaps rather than failures.
- Periodic reconciliation. Poll the provider's list endpoint on a schedule and compare against what you processed. This is the only method that catches silent 2xx, and it is the one most often skipped.
A daily reconciliation job that compares your record count against the provider's for the previous window will find classes of failure that no amount of endpoint monitoring will surface. It is also the cheapest thing on this list to build.
Recovery
Recovery depends entirely on whether you can replay. Before an incident, establish for each provider: is there a list or search endpoint that can return events for an arbitrary window, how far back does it reach, and is it rate limited in a way that makes a large backfill impractical? Record the answers. During an incident is the wrong time to discover that a provider's replay window is twenty-four hours.
Keeping that institutional knowledge attached to the integration rather than in an engineer's memory is exactly the kind of record Traxivo maintains, alongside the incident history that shows how often a given webhook has failed before, which is usually the more persuasive number when raising it with the provider.
A minimum viable webhook endpoint
- Verify the signature. Reject unverified payloads without processing.
- Persist the raw payload with its delivery identifier, before any parsing.
- Return 200 as soon as that write is durable.
- Process asynchronously, keyed idempotently on the provider's event identifier.
- Reconcile against the provider's own list endpoint on a schedule.
Four of those five steps are about being able to recover rather than about being correct on the happy path, which is the right ratio for anything delivered over an unreliable channel.
Provider differences worth recording
Webhook behaviour is not standardised, and the differences matter most during an incident. For each provider you consume, record the answers before you need them:
- Retry schedule and total window. Some give up within hours, others persist for more than a day. This determines how long an endpoint outage can last before loss is permanent.
- Replay or list endpoint. Whether events can be re-requested, how far back, and whether the rate limit makes a large backfill practical.
- Ordering guarantees. Almost always none, but a few providers offer sequence numbers that make gap detection trivial.
- Signature scheme and rotation. How secrets rotate, and whether old and new are both honoured during a window.
- Payload completeness. Whether the event carries the full resource or only an identifier you must then fetch, which changes your failure modes considerably.
That is five short answers per provider. Kept in the integration inventory, they convert most webhook incidents from an investigation into a lookup.
Frequently asked questions
What does a webhook delivery failure mean?
It means the sender considers an event delivered while the receiver never successfully processed it. The most common cause is an endpoint that returns a 2xx status before the event has been durably stored, so any later failure is invisible to the sender.
How should a webhook endpoint acknowledge an event?
Only after the raw payload is persisted somewhere it can be replayed from. Returning 200 should mean the event will not be lost, not that processing has completed. Parsing and business logic belong in an asynchronous step after the acknowledgement.
How do you detect webhooks that silently stop arriving?
Alert on the absence of expected volume rather than on errors, track gaps in any monotonic identifier the provider supplies, and run periodic reconciliation against the provider's list endpoint. Reconciliation is the only reliable way to catch events that were acknowledged but never processed.
Stop rediscovering the same integration failure
Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.
See how Traxivo works Browse use cases

