Chaos Engineering Platform  ·  View 09 of 21  ·  Data

Data Flow — Definition to Evidence

Six states a run's data passes through, and the two places it is deliberately thrown away.

Editable source SVG draw.io All views
Authored Definition Inventory snapshot Resolved Resolved targets Computed radius In flight Active leases Applied parameters High-res signals Judged Baseline window Verdict Retained Evidence bundle Downsampled series Audit entries Consumed Coverage report Findings Cost attribution selector resolve graph closure scoped to targets within caps agent reports collected compare held or refuted sealed after 90 days append aggregate if refuted per team Data Flow — From Definition to Evidence Data store Application we own synchronous event / async batch High-resolution signals exist only for evaluated signals, only in the baseline and injection windows. v 1.0 · owner Reliability Architecture · date 2026-09

Decisions

  • High-resolution signals are collected only for evaluated signals, and only inside the baseline and injection windows. This is the dominant cost lever in the whole design.
  • The applied fault parameters are recorded separately from the requested ones, because an adapter may clamp a value and a verdict read against the wrong parameters is worthless.
  • Coverage and cost marts are derived and rebuildable; the run record and audit log are not.

Assumptions

  • 500,000 samples/s during a large game day; ~100,000 run records a year with evidence bundles averaging 8 MB.
  • High-resolution series kept 90 days, then downsampled for the life of the run record.

Risks

  • Losing signal series inside the injection window degrades a verdict to INCONCLUSIVE. That is the intended behaviour and it is still a cost: the run has to be paid for twice.