API Gateway Platform  ·  View 18 of 21  ·  Operations

Observability

Eight signal families, one of which is the same data shown to the person it is about.

Editable source SVG draw.io All views
Emitted at the edge Buffered Stored Alerts on Who reads it Traffic (RED) rate / errors / duration in-process aggregation metric store, 13 mo SLO burn, 60 s SRE + API owner Fault attribution gateway vs upstream code tagged at source split series 5xx we caused SRE Policy decisions 429 by scope, 401 by cause never sampled analytics + audit denial spike per tenant Security + support Config state resident version per pod heartbeat coverage table split-brain > 60 s Rollout controller Dependency health counter, cache, upstream circuit state per-dependency series degraded mode entered SRE Traces W3C context, originated tail sample trace store, 7 days not alerted on Whoever is debugging Developer-facing per-app usage + errors same records 30-day window quota near ceiling The integrator Synthetic canary known call, every region out of band availability series silence is a failure SRE Observability — Signal Family by Pipeline Stage Row three is never sampled and row seven is the same data shown to the person it is about. The last row exists so a quiet dashboard cannot be mistaken for a healthy one. v 1.0 · owner Integration Platform Architecture · date 2026-09

Decisions

  • Fault attribution is its own signal family. A 502 and a 429 have different owners, and aggregating them into "errors" is the single most common way a gateway dashboard becomes useless during an incident.
  • Policy decisions are never sampled, at any volume — a denial nobody can investigate is not a control.
  • The synthetic canary exists so that a quiet dashboard cannot be mistaken for a healthy one.

Numbers

  • Alert on SLO breach within 60 s of onset; config split-brain alerts past 60 s (assumptions).
  • Access records sampled per route; errors always at full fidelity. Sampling rate is the main telemetry cost lever.

Risks

  • Developer-facing analytics and internal metrics are derived from the same records deliberately. If they diverge, a support conversation becomes an argument about whose numbers are right.