Customer 360 & Real-Time Risk Intelligence Platform  ·  View 17 of 20

Observability and SLA Monitoring

Signal types against pipeline stages — how each signal is emitted, collected, stored, turned into a detection, and acted on.

Editable source SVG draw.io All views
Emit
Emit
Collect
Collect
Store
Store
Detect
Detect
Act
Act
Metrics
Metrics
Broker, job, API
Broker, job, API
OTel collector
OTel collector
Time series store
Time series store
SLO burn alert
SLO burn alert
Scale or page
Scale or page
Logs
Logs
Structured JSON
Structured JSON
Node agents
Node agents
Log lake, 90 d
Log lake, 90 d
Error pattern rule
Error pattern rule
Ticket + runbook
Ticket + runbook
Traces
Traces
Spans on API & jobs
Spans on API & jobs
Tail sampling
Tail sampling
Trace store, 15 d
Trace store, 15 d
Latency outlier
Latency outlier
Hotspot fix
Hotspot fix
Pipeline SLA
Pipeline SLA
Lag, duration, offset
Lag, duration, offset
Airflow + Flink
Airflow + Flink
SLA history table
SLA history table
Miss prediction
Miss prediction
Escalate on call
Escalate on call
Data quality
Data quality
Expectation results
Expectation results
Quality event topic
Quality event topic
Quality history
Quality history
Threshold + drift
Threshold + drift
Quarantine + steward
Quarantine + steward
Cost
Cost
Job & cluster tags
Job & cluster tags
Billing export
Billing export
FinOps dataset
FinOps dataset
Budget variance
Budget variance
Rightsize, reserve
Rightsize, reserve
Observability and SLA Monitoring Matrix
Observability and SLA Monitoring Matrix
Application we own
Application we own
Interface / broker
Interface / broker
Data store
Data store
Decision point
Decision point
Queue / topic
Queue / topic
Data quality and cost are first-class signals, monitored on the same rails as infrastructure.
Data quality and cost are first-class signals, monitored on the same rails as infrastructure.
v 1.0 · owner Platform Engineering · date 2026-08
v 1.0 · owner Platform Engineering · date 2026-08
Text is not SVG - cannot display

First-class signals

  • Data quality and cost are monitored on the same rails as infrastructure metrics
  • Pipeline SLA is a measured signal, not an assumption drawn from job success
  • Every alert maps to a runbook; alerts without one are removed

SLO definitions

  • Freshness: 99% of enriched events within 2 seconds, measured hourly
  • Completeness: 99.9% of expected source records landed within the daily window
  • Availability: warehouse query success rate above 99.9% per rolling 30 days

Escalation

  • Consumer-lag and SLA-miss prediction page before a breach, not after
  • Quality breaches route to the owning steward, not to the platform on-call
  • Cost variance above budget threshold raises a FinOps review, not an incident