A real-time conferencing platform must serve unpredictable demand that once grew thirtyfold in ten weeks, where media relaying is bandwidth-intensive and latency-critical. Should it run on public cloud, its own data centres, or both - and what decides?
Show the full answer Hide the answer
The two workloads hiding inside one product
Media relaying is bandwidth-dominated, latency-critical, and — at the baseline — highly predictable. Egress pricing on public cloud makes sustained high-bandwidth workloads expensive in a way that does not improve with scale; it is the line item that grows fastest with success. Owning capacity changes the gradient of that curve rather than just its intercept.
Burst and geographic reach are the opposite: unpredictable, short-lived, and needed in places where building presence would take quarters. This is precisely what public cloud is good at, and paying a premium for capacity you use for six hours is a bargain compared to owning capacity you use for six hours.
Treating these as one workload forces a single answer to a question that has two.
Why the pure answers fail
Public cloud only. Egress and compute costs for sustained media relaying at scale become the dominant business expense, and the workload's baseline is stable enough that paying for elasticity you do not use is pure waste.
Owned only. Capacity acquisition has a lead time measured in months. When demand grows by an order of magnitude in weeks, the forecast is not merely wrong — it is unfixable within the window. There is no procurement process that responds to a step function.
Cheapest per unit today. The most seductive wrong answer. Unit cost ignores lead time, and lead time is the binding constraint during exactly the events that matter most.
The design that follows
- Baseline on owned capacity, sized to the predictable floor with a comfortable margin.
- Burst to cloud, with the capability continuously exercised rather than held in reserve. A burst path used for the first time during a surge is a burst path that does not work.
- Media routing that treats capacity as fungible, selecting a relay by health and latency rather than by ownership, so the mix can shift without a redesign.
- Geographic expansion via cloud first, converting to owned capacity only where sustained volume justifies it.
The constraint most teams under-weight
Lead time, not unit cost. The reason hybrid wins here is not arithmetic; it is that the two options have different response times to a demand change. Owned capacity is cheap and slow. Cloud capacity is expensive and immediate. A workload with a stable floor and violent peaks wants both, and the architecture's real job is to make the boundary between them a routing decision rather than a migration.
The general lesson
"Cloud or not" is almost never the right question. The right question is which workload, at which point in its demand curve — and the answer is frequently different for the baseline and the peak of the same system.