API Gateway Platform · View 18 of 21 · Operations
Decisions
- Fault attribution is its own signal family. A 502 and a 429 have different owners, and aggregating them into "errors" is the single most common way a gateway dashboard becomes useless during an incident.
- Policy decisions are never sampled, at any volume — a denial nobody can investigate is not a control.
- The synthetic canary exists so that a quiet dashboard cannot be mistaken for a healthy one.
Numbers
- Alert on SLO breach within 60 s of onset; config split-brain alerts past 60 s (assumptions).
- Access records sampled per route; errors always at full fidelity. Sampling rate is the main telemetry cost lever.
Risks
- Developer-facing analytics and internal metrics are derived from the same records deliberately. If they diverge, a support conversation becomes an argument about whose numbers are right.