Quiz
2667 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2667
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture78
Reliability & Resilience88
Observability81
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture91
Legacy Modernization92
AI-Era Architecture86
Software Architecture & Engineering84
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture90
Data Platform Architecture88
Streaming & Real-Time Data93
Data Governance & Semantics81
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture90
Assurance, Audit & Model Risk88
81 questions in Observability.
-
OpenTelemetry advanced
A platform team standardises every service on OTLP push to a collector and switches off Prometheus scraping. Metrics still arrive and dashboards still work. What has the team given up, and when does that bill arrive?
3 min answer opentelemetryprometheuspushscrape -
OpenTelemetry intermediate
A platform with services in five languages and three monitoring vendors considers adopting OpenTelemetry. What does it solve, and what is the realistic migration cost?
2 min answer opentelemetrystandardisationvendor-lock-inmigration -
OpenTelemetry intermediate
An organisation with several existing telemetry systems considers adopting OpenTelemetry. What does it actually solve, and what does adopting it not fix?
2 min answer segmentopentelemetrystandardsvendor-lock-in -
Profiling advanced
A data platform has good tracing and still cannot explain why a specific job is slow. What does tracing not tell you, and what does?
2 min answer databricksprofilingtracingcpu -
Profiling advanced Multiple choice
A developer-tools company needs to find a performance regression that only appears under real production workloads. What are the options for profiling in production, and what are their costs?
2 min answer profilingcontinuous-profilingsamplingoverhead -
Profiling advanced
A service is slow and CPU utilisation is 4%. What do you profile and what do you expect to find?
2 min answer profilingwall-clockblockingio -
Profiling advanced
Discord's 2020 post on rewriting its Read States service from Go to Rust described latency spikes on a roughly two-minute cadence, matching Go's forced garbage-collection interval, in a service that allocated very little. An engineer brings you a similar graph today and asks you to fund always-on profiling across 12000 containers. Walk me through what you would fund and what you would refuse.
3 min answer discordprofilinggarbage-collectiontail-latency -
SLO Monitoring advanced
A canary release looks healthy on p50 latency but a small set of enterprise tenants sees timeouts. Which metrics and gates should have caught it?
2 min answer canarytenant-segmentationtail-latencyrollout-gates -
SLO Monitoring intermediate Multiple choice
A checkout API has a 99.9% availability SLO and the team must decide where the indicator is computed from. The candidates are load-balancer access logs, in-process server metrics, the mobile client's own reporting, and synthetic probes. Which should be the primary source?
3 min answer sloslimeasurementavailability -
SLO Monitoring advanced
A communication platform sets a 99.9% availability SLO. How should alerting on that SLO be structured so it catches both sudden outages and slow degradation?
2 min answer sloburn-ratemulti-windowalerting -
SLO Monitoring intermediate Multiple choice
A food-delivery platform in Zomato's mould pushes order-status events to restaurant tablets and to customers through a queue-backed webhook fleet. The complaints are that status arrives late rather than that it never arrives. Which indicator should the SLO be written on?
3 min answer slofreshnessasynchronouswebhooks -
Structured Logging intermediate
A platform moves from free-text logs to structured logs. What becomes possible, and what discipline must accompany it to avoid making things worse?
2 min answer structured-loggingschemacardinalityquerying