Quiz
2667 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2667
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture78
Reliability & Resilience88
Observability81
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture91
Legacy Modernization92
AI-Era Architecture86
Software Architecture & Engineering84
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture90
Data Platform Architecture88
Streaming & Real-Time Data93
Data Governance & Semantics81
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture90
Assurance, Audit & Model Risk88
88 questions in Reliability & Resilience.
-
On-Call intermediate
A team is paged eleven times a week and morale is poor. What do you do first?
3 min answer oncallalert-fatiguesustainabilityarchitecture -
On-Call intermediate
An on-call rotation is producing burnout and slow responses. Alert volume is high and most pages are not actionable. What changes, and in what order?
2 min answer sliceoncallalertingactionability -
On-Call intermediate
You join a team of 12 engineers taking around 40 pages a week, of which perhaps 5 required action. The team is exhausted and two people have resigned. The director asks for a plan. What do you do, and in what order?
3 min answer on-callalert fatiguetoilsustainability -
Postmortems intermediate
A platform publishes detailed public postmortems for significant incidents. What does this practice cost, and what does it buy that internal postmortems do not?
2 min answer postmortemstransparencylearningtrust -
Postmortems intermediate
An engineer ran a command that deleted production data. What does the postmortem investigate?
2 min answer postmortemblamelesssystemicgoogle -
Postmortems advanced
An engineer runs a routine capacity-removal command with a typo. It removes far more than intended and a core service is down for hours. What does the postmortem conclude?
2 min answer awsblamelesstoolingblast-radius -
Postmortems advanced
An organisation writes thorough postmortems and keeps experiencing the same class of failure. What is missing?
2 min answer growwpostmortemaction-itemssystemic -
Redundancy advanced
A platform runs three redundant instances of a component and calculates its availability as extremely high. In practice, all three fail together during incidents. What is the flaw in the reasoning?
2 min answer redundancycorrelated-failureindependenceconfiguration -
Redundancy advanced
A platform runs three replicas of every service across three availability zones and still experiences total outages. What kinds of failure does that redundancy not address?
2 min answer zeptoredundancycorrelated-failureblast-radius -
Redundancy intermediate Multiple choice
Two service instances each run at 60% CPU behind a load balancer. Is that redundant?
2 min answer redundancycapacitycorrelationfailure-domains -
Capacity Planning advanced
A grocery delivery platform experiences a sudden multi-week increase in demand well beyond any forecast. Which capacity constraints bind first, and which cannot be solved with autoscaling?
2 min answer capacity-planningdemand-shockmarketplacesupply -
Capacity Planning intermediate
A service runs across three availability zones at 70% CPU during peak. The team wants to survive losing one zone at peak with no user impact. Roughly what does that require, and what does the same arithmetic say about running across two zones?
3 min answer capacityheadroomavailability zonesstatic stability