A partner needs order-status changes within a few seconds. Their estate is behind a corporate proxy that blocks inbound connections, their security team will not whitelist your webhook IPs, and their integration is a single long-running job server, not a browser. Which delivery channel fits?
Show the full answer Hide the answer
The deciding property
Who is able to open the connection, and in which direction the data needs to flow. The partner cannot accept inbound connections and the data flows one way, from you to them. That rules out half the options before latency, throughput or elegance are discussed, and it is the fact to extract first in any real integration conversation.
Why server-sent events
SSE is an ordinary GET that never ends, with the response body streamed as text/event-stream. The partner's job server dials out, which their proxy already permits, and reads events as they arrive. Three properties earn it here:
- It is HTTP all the way down. It passes proxies, authenticates with the same bearer token as the rest of your API, and needs no new network approval. That is the entire problem solved.
- Resumption is in the protocol. Each event carries an
id, and on reconnect the client sendsLast-Event-ID, so you replay from that point. You need a short retention buffer, but you do not need to design a resume mechanism. - The reconnect policy is standard. Clients back off and retry, and the server can tune the interval with a
retry:field.
The cost is real and bounded: one open connection and one server-side stream per subscriber, which means connection-count capacity planning rather than request-rate planning. A long-lived idle connection costs roughly tens of kilobytes of buffers plus whatever you hold per subscriber, so tens of thousands of partners is a fleet-sizing exercise, not a given.
Why not the others
- Webhooks are the default answer for partner event delivery and the right one most of the time: no connection state, no capacity coupling, and the partner scales their own receiver. Here the partner cannot host a reachable endpoint and will not whitelist your egress, so delivery has nowhere to land. Choosing webhooks anyway means weeks waiting on their network team.
- A WebSocket also dials out and would work, and it is the right call when the partner needs to send as well as receive, or needs binary frames. For one-way updates it buys a second protocol, a second set of proxy problems (some corporate proxies mishandle the upgrade), its own heartbeat and resume design, and a client library the partner must adopt. Bidirectional transport for unidirectional data is a cost with no matching benefit.
- Two-second polling with a cursor is the honest fallback and is better than people admit: no connection state, trivially debuggable, and it survives any network. Here it means 1,800 requests per hour per partner, almost all returning empty, and a worst-case staleness of two seconds plus processing. For one partner that is fine. As a pattern across hundreds of partners it is a large permanent load, and the natural response — raising the interval — breaks the latency requirement.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| The partner can host an endpoint and whitelist egress | Webhooks | No connection state and they own scaling |
| The partner must send commands as well as receive | WebSocket | Duplex is the actual requirement |
| Acceptable staleness rises to a minute or more | Cursor polling | Simplest thing that works and no long-lived state |
| Subscribers grow past tens of thousands | Webhooks or polling | Connection count becomes the binding constraint |
Build the cursor-polling endpoint regardless. Every streaming channel needs a catch-up path for a subscriber that was down longer than your buffer, and that path is the polling endpoint you were going to skip.
When not to stream at all
If the partner's actual requirement is "the same day", streaming is the wrong shape and a scheduled file drop is better: no connection state, no retention buffer, and nothing to page anyone about at 03:00. Weak answers choose the channel from how modern it sounds rather than from who can open the socket, and then spend a month on a security exception they did not need.