Chaos Engineering Platform  ·  View 14 of 21  ·  Runtime

Guardrail Evaluation

Seven independent gates, each evaluated at four points in a run's life.

Editable source SVG draw.io All views
Pre-flight At issue During injection At abort Blast radius Resolve selector Compute 3 measures Reserve radius Re-check on scale-up Release reservation Dependency closure Graph freshness ≤ 24 h Downstream set Refuse shared SPOF Watch collateral SLIs Record reached set Error budget Read SLO budget Debit chaos share Stop at exhaustion Settle actual spend Windows & incidents Check exclusion window Confirm no open incident Subscribe to incident feed Skip, do not queue Authorisation Ownership check Tier approval Sign lease to target Renew only while valid Audit the whole chain Reversibility Adapter has revert Proven in non-prod ≤ 30 d Scope to one class Report applied state Verify reversion Telemetry coverage SLIs exist and are fresh Baseline established Completeness threshold Mark INCONCLUSIVE Guardrail Evaluation — Per Check Class, Per Phase Application we own Security / platform Risk / gap Data store External / third party Interface / broker Decision point Seven independent gates. Any one refusing is a refused run; none of them can be waived by the requester. v 1.0 · owner Reliability Architecture · date 2026-09

Decisions

  • Seven gates, none waivable by the requester. An over-cap run needs a named approver; it does not get to skip the check.
  • Blast radius is re-checked during injection, because a scale-up event can turn a compliant 5% into a non-compliant one without anybody changing the definition.
  • Reversibility is a pre-flight gate: an adapter without a proven revert, exercised in non-production within 30 days, is ineligible for production.

Assumptions

  • Default caps: ≤ 5% of healthy replicas, ≤ 1 zone, ≤ 1% of production request volume.
  • Chaos error budget ≤ 10% of a service's monthly SLO error budget.
  • Dependency graph freshness ≤ 24 hours, enforced as a block.

Risks

  • An over-conservative dependency closure can block so many runs that teams route around the platform. That failure is quieter than an outage and worse for the practice.