LinkedIn Professional Network  ·  View 26 of 30  ·  6 · Operations

Graceful Degradation Contract

What each capability does when a dependency is slow, down, or in a lost colo, agreed before the incident.

Editable source SVG draw.io All views
Healthy Dependency slow Dependency down Colo lost Feed Ranked, personalised Skip recs, keep network Chronological + trending Served by next colo Search Federated, personalised Fewer verticals Search off, rest up Served by next colo Notifications Seconds to device Queued, delayed Backlog; actions succeed Replayed from Kafka Messaging Under 500 ms Push, not SSE Store and forward Reconnect elsewhere Profile read Cache + Espresso Stale cache allowed Espresso direct Served by next colo Job apply Strongly consistent Retry with same key Fail closed, retry Idempotent replay Graceful Degradation Contract Application we own Opportunity Risk / gap Each row degrades on its own non-critical dependency; red cells are where a member notices. v 1.0 · owner Site Reliability Engineering · date 2026-09

Decisions

  • Every row has an owner and a flag. The degraded mode is chosen in advance, not during an incident
  • Core writes (post, apply, message, connect) never depend on search, recommendations or notifications
  • The requirement's three examples map directly: recommendations down gives a chronological feed; notifications down, actions still succeed; search down, profile, feed and apply stay up

Mechanisms

  • Timeouts and circuit breakers per dependency, in the Rest.li client
  • Hodor sheds bot and prefetch traffic first
  • Kafka holds a notification backlog for up to 7 days

The red cells

  • Search switched off is visible to members, and accepted
  • Apply fails closed, because a lost application is worse than a retry