LinkedIn Professional Network  ·  View 22 of 30  ·  6 · Operations

Deployment Topology

Four active colos, members pinned to one of them, and a colo drained as routine rather than as an emergency.

Editable source SVG draw.io All views
Global edge Global DNS health-checked Azure Front Door 165+ PoPs · WAF Media CDN cache · signed URLs Colo A · US · one of four serving colos Fault zone 1 Traffic tier ATS Rest.li services stateless pool Espresso replicas partition leaders Kafka brokers rack-aware Fault zone 2 Traffic tier ATS Rest.li services stateless pool Espresso replicas followers Kafka brokers rack-aware Fault zone 3 Traffic tier ATS Rest.li services stateless pool Espresso replicas followers Kafka brokers rack-aware Colos B, C and D · identical stack, all serving Colo B · US Full stack active Colo C · US Full stack active Colo D · APAC Full stack active Offline grid HDFS · Spark · Azkaban pinned colo in-colo replica Brooklin mirror LinkedIn — Deployment Topology Interface / broker Application we own Data store Queue / topic synchronous event / async Fault zones inside a colo are this design's assumption; LinkedIn has published colo count, not intra-colo layout. v 1.0 · owner Site Reliability Engineering · date 2026-09

Decisions

  • Four serving colos, all active (LinkedIn, 2017). The traffic tier pins each member to one colo with a signed cookie
  • Losing a colo is routine: TrafficShift drains it by re-pinning members elsewhere, the same mechanism LinkedIn uses for load tests (LinkedIn, 2017)
  • The global edge has been on Azure Front Door since 2020, replacing LinkedIn's own 19 PoPs

RPO and RTO

  • Service failure: restart in seconds, no data loss
  • Fault-zone loss: RPO 0 through in-colo replicas; RTO under 1 minute, automatic
  • Colo loss: RPO of seconds through async cross-colo replication; RTO under 15 minutes by traffic shift. Derived stores are rebuilt, so RPO does not apply to them

Risks

  • A write acknowledged in a colo that then dies can be lost for seconds. Credentials and applications reconcile after failover through a written playbook
  • Disaster recovery works only if it is exercised, so colos are drained on a schedule