Practice this topic in a realistic system design interview
Suppose your application uses a payment provider such as Stripe. When a payment succeeds, your backend needs to update the order status.
One option is to repeatedly ask the payment provider whether the payment has completed.
Polling wastes work when most checks return "nothing changed." It also adds delay, because your system learns about the payment only on the next check.
The other option is to register a webhook endpoint. When the payment succeeds, the payment provider sends an HTTP request to your server containing the event.
This chapter covers what a webhook is, how delivery works, how to make it reliable with retries, idempotency, and queues, how to handle ordering and security, and how to design the sending side.
A webhook is an HTTP callback. One system exposes an HTTP endpoint, and another system sends an HTTP request to that endpoint whenever a particular event occurs.
The sending system is usually called the provider. The receiving system exposes the endpoint and is called the consumer or receiver.
In the payment example, Stripe is the provider and your backend is the consumer. Instead of your system asking for updates, the provider tells you when something happens.
Loading simulation...
A webhook request usually contains:
payment_intent.succeededA typical body looks like this:
Delivery details and security information usually travel in headers. The exact names vary by provider:
| Header | Purpose |
|---|---|
Content-Type | Usually application/json |
| Event type header | Event name, such as pull_request or payment_intent.succeeded |
| Delivery ID | Unique ID for this delivery attempt or event |
| Signature | HMAC or similar signature used to verify who sent the request |
| Timestamp | Helps detect replayed requests when used with a signature |
| User agent | Identifies the sender, such as GitHub or Stripe |
Your server receives the request, verifies it, processes the event, and returns a successful HTTP response.
In production systems, sending the request is the easy part. The hard part is making delivery reliable. A webhook is still an HTTP request crossing the internet. It can fail, be sent more than once, or arrive out of order. The rest of this chapter deals with those problems.
Suppose a provider sends an event, but your server is temporarily unavailable. The provider should not discard the event. Instead, it stores it and retries later, usually with exponential backoff. For example, it might retry after 1 minute, then 2, 4, and 8 minutes.
Providers also add jitter, a small random offset on each delay, so thousands of failed deliveries do not all retry at the same moment.
Eventually, retries must stop. Most systems define a maximum number of attempts or a retry window, such as 3 days. After that, the event moves to a failed-delivery queue, where someone can inspect it or replay it manually.
Retries introduce another problem: duplicate delivery.
Suppose your server processes a webhook successfully, but the network fails before the provider receives the HTTP response. Your server knows the event succeeded. The provider does not. From its side, the safest option is to retry, so your application receives the same event again.
That is why most webhook systems provide at-least-once delivery, not exactly-once delivery. Your consumer must handle duplicate events safely.
The standard solution to duplicate delivery is idempotency: processing the same event twice has the same effect as processing it once.
Every webhook event should have a unique event ID. When your application receives an event:
Store processed IDs in reliable storage with a unique constraint, so two concurrent deliveries of the same event cannot both slip through:
When a duplicate arrives, return success, because the original was already accepted:
This does not create true exactly-once delivery. It makes repeated deliveries safe, which is the practical goal.
Another common mistake is doing too much work inside the webhook request itself.
Imagine your server receives a payment webhook. The handler updates the database, sends an email, generates an invoice, updates analytics, and calls another service, all before responding. If that takes too long, the provider hits its timeout and assumes delivery failed. Then it retries, and you might process the event again while the first request is still running.
A more reliable pattern keeps the webhook endpoint lightweight:
Background workers then handle the expensive processing asynchronously. This separates webhook delivery from business logic, so slow downstream systems cannot delay the acknowledgment.
The order matters. Do not return 200 OK before the event is safely stored. If your process crashes after returning success but before saving, the provider has no reason to retry, and the event is lost.
At scale, a common webhook consumer architecture looks like this. The endpoint receives the request, validates it, writes the event to a durable queue, and immediately returns success. A pool of workers consumes events from the queue and performs the actual business logic.
This gives three advantages:
The endpoint stays fast, while the queue acts as a buffer between external traffic and internal processing.
One detail to get right: if the endpoint saves the event to a database and then sends it to a separate queue as two independent writes, a crash between them can drop the event. Either commit both together, often with an outbox table that a relay forwards to the queue, or let workers read directly from the events table.
Another subtle problem is event ordering.
Suppose a system sends two events. The first says an order was created. The second says the order was cancelled. You might assume they arrive in that order, but distributed systems rarely give you that guarantee. The first webhook could fail and be retried later, while the second succeeds immediately. Your application receives the cancellation before the creation.
There are two common solutions:
For payments, the second approach is common: before fulfilling an order, confirm its status through the provider's API.
Webhook endpoints are publicly accessible. Without verification, anyone who finds your endpoint could send a fake payment_succeeded event and get free access.
So your server needs a way to verify that a request actually came from the trusted provider. The common approach is request signing:
Two details matter:
In real code, use the provider's official library when one exists, because signature formats differ between providers.
A webhook endpoint should stay simple. Verify the request, validate the event, store it durably, and return a response quickly. Avoid running complex business logic directly inside the request.
HTTP status codes matter. A success response tells the sender the event was accepted. A server error usually signals that it should retry. A permanent client error may indicate that retrying will not help.
| Response | Meaning |
|---|---|
200 OK or 204 No Content | Event was accepted or already handled |
400 Bad Request | Request body is malformed or unsupported |
401 Unauthorized / 403 Forbidden | Signature or authorization check failed |
404 Not Found | Endpoint is wrong or no longer exists |
429 Too Many Requests | Receiver is overloaded; treated as a retryable failure by most providers |
500 / 503 | Temporary receiver failure; provider should retry |
The exact behavior depends on the provider, so retry semantics should be clearly defined on both sides.
Observability is just as important. Webhook failures are easy to miss: the provider retries quietly, and users notice only the symptom, such as an order stuck in PENDING. Track:
Without that visibility, webhook issues are hard to detect and debug.
So far, the focus has been on consuming webhooks. Now look at the other side: sending them.
Suppose your platform delivers webhooks to thousands of customers. When an event occurs, the service that created it should usually not send the webhook directly. Instead:
Separating event generation from webhook delivery makes the system more reliable and easier to scale. The service that created the event never waits on a customer's server.
If you operate a webhook provider, some customer endpoints will be slow, unreliable, or completely down.
One broken destination should not affect everyone else. So isolate deliveries per endpoint, each with its own:
You may also need concurrency limits, so you do not overwhelm a customer with too many requests at once. This matters most when replaying a large backlog after an outage: a customer that was down for an hour should receive the backlog at a pace it can handle, not all at once.
Webhooks are a strong choice when one system needs to notify another that an event occurred. Common examples:
They work especially well for discrete events, where a permanent connection is unnecessary.
A webhook is an HTTP callback: when an event occurs, the provider sends an HTTP request to an endpoint you registered, instead of you polling for it.
Sending the request is easy. Reliable delivery is the hard part. Providers retry with exponential backoff and jitter, then move failed events to a failed-delivery queue. Retries mean duplicates, so consumers use event IDs for idempotency. Events can arrive out of order, so use version numbers or fetch the latest state.
A good endpoint verifies the signature over the raw body, stores the event durably, returns quickly, and leaves the real work to queue-backed workers. A good provider separates event generation from delivery and isolates each customer endpoint.
Reliable webhooks need more than an HTTP request: durable event storage, retries, idempotency, authentication, monitoring, and a strategy for failed deliveries.
12 quizzes