Chaos Engineering Platform  ·  View 01 of 21  ·  Context and scope

System Context

Who uses the platform, which systems it only reads, and what it is allowed to touch.

Editable source SVG draw.io All views
People Service owner Reliability engineer Tier-1 approver On-call responder Systems it reads Service catalogue Dependency graph SLO store Cloud Monitoring Identity provider Chaos Engineering Platform Bounded failure, on purpose Systems it acts on and tells Production service fleet 900 services Incident platform Deployment pipeline runs experiments runs game days approves notified ownership graph edges SLI defs signals OIDC faults injected findings verdict Chaos Engineering Platform — System Context Person or role External / third party Security / platform Application we own synchronous Out of scope: load testing, security red-teaming, functional test automation. v 1.0 · owner Reliability Architecture · date 2026-09

Decisions

  • The service catalogue, dependency graph and SLO store are read-only inputs. The platform never becomes a second place where ownership or SLIs are defined.
  • Steady state is evaluated against the target's own production SLIs — the ones its owners are paged on — so a passing experiment means something to the people who own the service.
  • The customer does not appear on this view. A run a customer notices has already failed its guardrails.

Assumptions

  • 900 services, 40,000 pods, six regional clusters in three regions — a stated assumption for sizing, not a measurement.
  • The incident platform can be queried for open incidents of a given severity, and can accept a finding as a tracked item.

Out of scope

  • Load and performance testing, security red-teaming, and functional test automation.
  • Owning the SLO definitions, the service inventory, or the on-call rotation.