pattern

Event-Driven Integration

Integrating systems by delivering events rather than by polling or synchronous calls — including the delivery guarantees the receiver must be told about.

webhookseventsintegrationshopifydelivery

Definition

Rather than a consumer polling for changes or the producer calling the consumer synchronously, the producer emits events that consumers receive asynchronously. Externally this usually takes the form of webhooks; internally, a message broker or event log.

Why webhooks are harder than they look

The producer is now making an outbound HTTP call to an endpoint it does not control, which may be slow, down, wrong, or malicious. Every property of a resilient client applies, in reverse:

  • Delivery is at-least-once. Network failures mean retries, and retries mean duplicates. The receiver must deduplicate on an event ID, and the documentation must say so loudly.
  • Ordering is not guaranteed. A refunded event can arrive before the succeeded event it refers to. Consumers should fetch current state rather than assume the event sequence tells the whole story.
  • Retries need a schedule and an end. Exponential backoff over hours or days, then a dead-letter state visible to the integrator, plus a way to replay.
  • Signing is mandatory. The receiver must be able to verify the event came from you and has not been altered, using a shared secret and a timestamp to prevent replay.
  • Slow consumers must not affect the producer. Delivery runs in a bounded worker pool with per-endpoint circuit breaking, or one slow integrator degrades delivery for everyone.

Industry example

Commerce platforms live on this pattern: a merchant's order, inventory and fulfilment systems all integrate through webhooks, and the platform has no control over any of them. Shopify's approach illustrates the necessary defensive posture — events are signed, delivery is retried on a defined schedule, endpoints that fail persistently are disabled and the merchant is notified, and an API exists to reconcile anything missed.

That last mechanism is the one most often omitted and the most important. Webhooks are an optimisation over polling, not a guarantee. A robust integration reconciles periodically against the API, because some events will be missed — the endpoint was down for a day, a deployment lost them, a bug dropped them silently. A design that treats webhook delivery as authoritative will eventually diverge with no way to detect it.

Failure scenarios

  • Consumer processes synchronously inside the webhook handler, times out, and the producer retries — creating duplicate work and eventual endpoint disabling. Receivers should acknowledge immediately and process asynchronously.
  • No signature verification, so anyone who learns the URL can inject events.
  • Ordering assumed, producing state machines that break on out-of-order arrival.
  • No reconciliation, so silent divergence accumulates.
  • A retry storm when an integrator recovers and thousands of queued events arrive at once.

Trade-offs

Bought: near-real-time integration without polling load, decoupling, and extensibility. Sold: exactly the delivery guarantees people assume they have, plus an operational surface — delivery queues, per-endpoint health, dead-letter handling — that the producer must own on behalf of consumers.

Interview question

"Design webhook delivery for a platform with 50,000 integrators. What happens when one endpoint is down for two days?"