Quiz
2667 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2667
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture78
Reliability & Resilience88
Observability81
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture91
Legacy Modernization92
AI-Era Architecture86
Software Architecture & Engineering84
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture90
Data Platform Architecture88
Streaming & Real-Time Data93
Data Governance & Semantics81
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture90
Assurance, Audit & Model Risk88
53 questions in Reliability & Resilience.
-
Reliability Culture advanced
An airline's crew-scheduling system cannot re-plan fast enough during cascading weather disruption, forcing manual processes that cannot keep up. What signals indicated this risk earlier, and how should such a system be modernised?
3 min answer southwestlegacytechnical-debtcapacity -
Reliability & Resilience advanced
A super-app combines messaging, payments, social feeds, mini-programs and notifications in one product used by hundreds of millions daily. What is the dominant reliability risk, and what structural property addresses it?
2 min answer reliabilityisolationsuper-appblast-radius -
Reliability & Resilience advanced
A workflow executes over days or weeks while individual workers restart many times. How should workflow state, timers, retries, heartbeats and task ownership be designed?
2 min answer temporaldurable-executionworkflowstimers -
Reliability & Resilience advanced
In July 2019 a single regex deployed globally took Cloudflare's network to 100% CPU within seconds. What does this say about how configuration should be released?
2 min answer deploymentblast-radiusconfigcase-study -
Reliability & Resilience advanced
Maersk rebuilt roughly 4,000 servers and 45,000 PCs in about ten days after NotPetya in 2017, and recovered its directory only because one data centre had been offline during the attack. What does this say about DR design?
2 min answer disaster-recoveryransomwarebackupcase-study -
Resilience Testing advanced
A brief database slowdown caused a two-hour full outage. Explain the likely amplification chain and the fixes at each stage.
2 min answer cascading-failureretriestimeoutsincident -
Resilience Testing advanced
A travel platform depends on hundreds of external suppliers with varying reliability. How should resilience testing be designed when the failures originate outside the system?
2 min answer resilience-testingfault-injectionexternal-dependenciesexpedia -
Resilience Testing advanced
How would you test that your service degrades correctly when a dependency's latency rises tenfold?
2 min answer resilience-testinglatency-injectiontimeoutslittle's-law -
Resilience Testing advanced
Review this resilience programme. Chaos experiments run weekly in staging at 03:00, they inject only instance termination, results are recorded in a spreadsheet, and there is a kill switch that has never been used. What would you change, and what would you keep?
3 min answer chaos engineeringgame daysstagingsteady state -
RTO & RPO advanced
A collaborative workspace product must define RTO and RPO. The product team says "we can never lose a user's work". What does that requirement actually mean, and what does it cost?
2 min answer rportodurabilitycollaboration -
RTO & RPO advanced
A stakeholder asks for zero RPO across the estate. What does that cost and what would you propose instead?
2 min answer rporeplicationcostbusiness-continuity -
RTO & RPO advanced
In April 2022 a maintenance script at Atlassian used the wrong identifiers and deleted sites belonging to about 775 customers. Restoration took up to about two weeks for some of them, even though backups existed and were working. What property of the recovery design accounts for that gap?
3 min answer atlassianrestore granularitymulti-tenantblast radius