Quiz
2827 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2827
Architecture Fundamentals91
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture87
Reliability & Resilience99
Observability92
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture100
Legacy Modernization92
AI-Era Architecture96
Software Architecture & Engineering93
Architecture Patterns84
Architecture Decision-Making101
The Architect's Meta-Skills92
Delivery & Release Engineering103
Platform Engineering & Developer Experience92
Testing & Quality Architecture102
Data Platform Architecture98
Streaming & Real-Time Data93
Data Governance & Semantics91
Frontend & Experience Architecture101
Edge, Mobile & IoT98
Regulatory & Data Protection Architecture99
Assurance, Audit & Model Risk98
2827 questions.
-
Failure Modes advanced
One instance in a fleet of 20 is failing 30% of its requests. Health checks pass, aggregate error rate is 1.5%, no alerts fire. How do you detect and handle this?
2 min answer gray-failureobservabilityresilience -
Failure Modes advanced
One instance in a fleet of fifty is returning correct responses very slowly. Health checks pass and it stays in rotation. How do you detect and handle this?
2 min answer grey-failurehealth-checksdetectionmitigation -
Failure Modes advanced
Roblox's 2021 outage lasted 73 hours after a service-discovery and key-value cluster degraded under contention. Trace the failure chain, and identify the three architectural properties that turned a degradation into a three-day outage.
2 min answer robloxconsulservice-discoverycascading-failure -
Failure Modes advanced
What happens when a globally distributed real-time communication service loses an entire region - and how should DNS, health checks, traffic steering, failover, connection draining and data consistency interact?
3 min answer failure-modesregional-failurefailoverdns -
Failure Thinking intermediate
A major platform launch is three months away. How would you run a pre-mortem, and why bother?
2 min answer riskfacilitationmeta-skills -
Failure Thinking advanced
Between August and early September 2025 three separate infrastructure bugs degraded Claude's output quality - one affected 16% of Sonnet 4 requests in the worst hour of 31 August - while error rates and latency stayed normal. What class of failure is this, and what would you have had to build beforehand to notice it?
2 min answer anthropicsilent-failuredetectionquality-regression -
Failure Thinking advanced
What does it mean to design by thinking about failure first, and what does that produce that requirement-driven design does not?
2 min answer credfailure-modesdesignpremortem -
Failure Thinking advanced
What does it mean to think in failure modes as a habit rather than as a checklist item, and which questions consistently produce findings?
3 min answer failure-modesdesign-reviewresiliencequestions -
Failure Thinking advanced
What questions characterise failure thinking, and which are most often skipped in design reviews?
2 min answer failure-thinkingreviewdependenciesdegradation -
Fault Isolation intermediate
A design platform's synchronous editing experience and its asynchronous export rendering share a compute cluster. Exports occasionally saturate it and the editor becomes unusable. What would you change?
2 min answer fault-isolationbulkheadsworkload-separationcanva -
Fault Isolation advanced
A financial platform's card authorisation path must survive dependency failures, deployments, overloaded downstreams and partial network failures. How should isolation, deadlines, breakers, bulkheads, caching, shedding and fallback combine?
2 min answer brexrampauthorisationisolation -
Fault Isolation advanced
A security vendor pushes a content update to millions of endpoints simultaneously and a malformed file crashes them all at kernel level. What should have been in place, and why is "it was data, not code" the wrong defence?
3 min answer crowdstrikestaged-rolloutblast-radiuskernel