Event-Driven Notification Platform · View 21 of 26 · 5 · Operations
Observability and Operations
Every signal, where it goes, and how a support engineer answers the question in one place.
Copy
PNG
PDF
⋯
Editable source
SVG
draw.io
All views
Emit
Emit
Collect
Collect
Store
Store
Consume
Consume
Act
Act
Metrics
Metrics
OpenTelemetry SDK
RED per service
OpenTelemetry SDK...
OTel Collector
OTel Collector
Prometheus and Mimir
13 months
Prometheus and Mimir...
SLO dashboard
Grafana
SLO dashboard...
Page on burn rate
Alertmanager
Page on burn rate...
Logs
Logs
Structured JSON
event_id · correlation_id
Structured JSON...
Fluentd
PII redaction filter
Fluentd...
Loki
30 d
Loki...
Incident search
Incident search
Runbook link
Runbook link
Traces
Traces
W3C tracecontext
propagated on the bus
W3C tracecontext...
OTel Collector
tail sampling · errors kept
OTel Collector...
Tempo
7 d
Tempo...
Event to provider span
one trace end to end
Event to provider span...
Latency budget alert
Latency budget alert
Notification lifecycle
Notification lifecycle
State transitions
CREATED to DELIVERED
State transitions...
Receipt topic
Receipt topic
ClickHouse
90 d attempts
ClickHouse...
Notification timeline
one query per recipient
Notification timeline...
Support answer
target under 2 min
Support answer...
Queue and backlog
Queue and backlog
Consumer lag
per tier per tenant
Consumer lag...
Kafka exporter
Kafka exporter
Prometheus
Prometheus
Backlog board
DLQ depth included
Backlog board...
KEDA scale-out
before the page
KEDA scale-out...
Tenant and cost
Tenant and cost
Per-tenant counters
sent · failed · suppressed
Per-tenant counters...
OTel Collector
OTel Collector
ClickHouse
ClickHouse
Chargeback report
provider spend per tenant
Chargeback report...
Quota enforcement
Quota enforcement
Observability — Signals, Storage and Who Gets Woken
Observability — Signals, Storage and Who Gets Woken
Application we own
Application we own
Interface / broker
Interface / broker
Data store
Data store
Security / platform
Security / platform
Queue / topic
Queue / topic
The fourth row is the one the requirement asks for by name. Because event_id, notification_id and provider_ref are on the same trace and the same ClickHouse row, why customer X did not receive notification Y is one query, not four systems.
The fourth row is the one the requirement asks for by name. Because event_id, notification_id and provider_ref are on the same trace and the same ClickHouse row, why customer X did not receive notification Y is one query, not four systems.
v 1.0 · owner Data & AI Global Practice · date 2026-08
v 1.0 · owner Data & AI Global Practice · date 2026-08
Text is not SVG - cannot display
The question this has to answer
Why did customer X not receive notification Y — answerable from one query, target under two minutes
It works because event_id, notification_id, provider_ref and the suppression reason are on the same trace and the same ClickHouse row
A suppressed notification is a first-class record with a reason, so the answer is never an absence of data
Signals and cost control
Tail sampling keeps every error trace and 1 percent of successes, which is what makes tracing affordable at this volume
Logs are redacted at the collector, so PII cannot reach Loki even if a service logs it
Per-tenant counters drive both the chargeback report and quota enforcement from the same series
Alerting
Pages are on SLO burn rate and on P0 DLQ depth above zero; everything else is a ticket
KEDA scales out on consumer lag before the lag alert fires, so a page means scaling did not help
Every alert carries a runbook link; an alert without one fails the release gate
◀ Autoscaling, Capacity and Backpressure
All views
Administration and Operations Surface ▶