Observability Platform  ·  View 01 of 25  ·  Context and scope

System Context

Who reads it, what it observes, what it hands off — and the three neighbouring systems it is often mistaken for.

Editable source SVG draw.io All views
What it observes 900 in-house services EKS · EC2 · managed AWS Web & mobile clients People who read it On-call engineer Service owner SRE / platform Security engineer Systems it hands off to Paging platform Service catalogue Identity provider Observability Platform metrics · logs · traces Deliberately outside the boundary Regulatory audit log Product analytics Incident management page to cause SLOs, budgets estate health compliance OTLP infra metrics untrusted path firing alert ownership identity own custody not here handover Observability Platform — System Context Application we own Security / platform External / third party Person or role synchronous event / async Routing, escalation and acknowledgement belong to the paging platform. This set stops at the firing alert. v 1.0 · owner Reliability Architecture · date 2026-09

Decisions

  • The platform stops at the firing alert. Routing, escalation, on-call schedules and acknowledgement belong to the paging platform, which is a separate product with a separate availability argument.
  • Product analytics and the regulatory audit log are named as out of scope on the diagram rather than left unmentioned, because both get pushed into an observability platform by default and both make an operational store expensive and slow.
  • Untrusted client telemetry (web and mobile) enters through its own path and is drawn as an external producer, not as another service.

Assumptions

  • 900 in-house services, 12,000 hosts, 40,000 containers, three AWS regions, 70 owning teams, ~450 engineers.
  • A service catalogue already exists and is authoritative for ownership. Without it, the tenancy model in view 08 has no source of truth and the platform gains a manual onboarding step it cannot scale.

Risks

  • The platform runs on the estate it observes. That circular dependency is answered in views 19 and 21 and nowhere else; if the self-telemetry account is descoped for cost, the platform can be blind and report itself healthy.
  • Four reader groups with genuinely different questions is the usual way an observability console becomes unusable for all of them.