Quiz
2667 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2667
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture78
Reliability & Resilience88
Observability81
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture91
Legacy Modernization92
AI-Era Architecture86
Software Architecture & Engineering84
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture90
Data Platform Architecture88
Streaming & Real-Time Data93
Data Governance & Semantics81
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture90
Assurance, Audit & Model Risk88
42 questions in Observability.
-
Observability advanced
An error-tracking platform receives millions of events, many of them the same underlying problem with different stack details. How should grouping, deduplication, high-cardinality metadata and retention be designed?
2 min answer sentrygroupingdeduplicationretention -
Observability advanced
On 11 December 2024 OpenAI rolled a new telemetry service out across every Kubernetes cluster. Within about half an hour the API servers were saturated and services could no longer resolve one another; full recovery took until the evening. What turned a monitoring change into a total outage, and which of the contributing factors would you fix first?
3 min answer openaikubernetescontrol planedns -
Observability advanced
Your observability bill is now 40% of your compute bill. Leadership wants it cut in half without going blind. What do you cut?
2 min answer observabilitycostcardinalitysampling -
Observability advanced
p99 latency on checkout tripled overnight. Dashboards look normal, no deployment went out, and every service reports healthy. How do you find it?
3 min answer debuggingtracinglatencydiagnosis -
OpenTelemetry advanced
A fleet exports OTLP to a gateway collector deployment with memory_limiter first in the pipeline, then a batch processor, then an exporter with a sending queue. The telemetry backend starts answering in 8 seconds instead of 80 ms, and request volume doubles at the same time. What happens second by second, and what stops it?
2 min answer opentelemetrycollectorbackpressureload-shedding -
OpenTelemetry advanced
A platform team standardises every service on OTLP push to a collector and switches off Prometheus scraping. Metrics still arrive and dashboards still work. What has the team given up, and when does that bill arrive?
3 min answer opentelemetryprometheuspushscrape -
Profiling advanced
A data platform has good tracing and still cannot explain why a specific job is slow. What does tracing not tell you, and what does?
2 min answer databricksprofilingtracingcpu -
Profiling advanced Multiple choice
A developer-tools company needs to find a performance regression that only appears under real production workloads. What are the options for profiling in production, and what are their costs?
2 min answer profilingcontinuous-profilingsamplingoverhead -
Profiling advanced
A service is slow and CPU utilisation is 4%. What do you profile and what do you expect to find?
2 min answer profilingwall-clockblockingio -
Profiling advanced
Discord's 2020 post on rewriting its Read States service from Go to Rust described latency spikes on a roughly two-minute cadence, matching Go's forced garbage-collection interval, in a service that allocated very little. An engineer brings you a similar graph today and asks you to fund always-on profiling across 12000 containers. Walk me through what you would fund and what you would refuse.
3 min answer discordprofilinggarbage-collectiontail-latency -
SLO Monitoring advanced
A canary release looks healthy on p50 latency but a small set of enterprise tenants sees timeouts. Which metrics and gates should have caught it?
2 min answer canarytenant-segmentationtail-latencyrollout-gates -
SLO Monitoring advanced
A communication platform sets a 99.9% availability SLO. How should alerting on that SLO be structured so it catches both sudden outages and slow degradation?
2 min answer sloburn-ratemulti-windowalerting