Quiz
2627 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2627
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture77
Networking86
API & Integration Architecture78
Reliability & Resilience88
Observability81
Performance & Capacity Engineering90
Security Architecture85
Cost Architecture & FinOps82
Business Architecture93
Architecture Communication91
Enterprise Architecture91
Legacy Modernization82
AI-Era Architecture86
Software Architecture & Engineering84
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture90
Data Platform Architecture88
Streaming & Real-Time Data93
Data Governance & Semantics81
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture90
Assurance, Audit & Model Risk88
88 questions in Reliability & Resilience.
-
DR Testing advanced
An auditor requires evidence that a four-hour RTO is achievable for a system that has never failed over. Production cannot be risked and the business will not accept an unplanned outage. Sequence the first real disaster-recovery test.
3 min answer dr testingrtorehearsalevidence -
DR Testing intermediate
An enterprise runs an annual disaster recovery test that always succeeds, yet the team has low confidence in real recovery. What is likely wrong with the test?
2 min answer dr-testingdrillsrealismenterprise -
DR Testing advanced
In January 2017 GitLab lost roughly six hours of database data after an engineer deleted a directory on the wrong host during replication troubleshooting, and then found that several backup and replication mechanisms had silently not been working. What signal would have revealed the broken backups beforehand, and why did nothing report them?
3 min answer gitlabbackupsrestore testingsilent failure -
DR Testing advanced
Your DR plan is pilot light with a 30-minute RTO. What would you test, and what will the test probably reveal?
2 min answer drtestingcontrol-plane -
Error Budgets advanced
A platform has an error budget policy stating that feature work stops when the budget is exhausted. The budget is exhausted, and product leadership wants a major launch to proceed. How should this be resolved?
2 min answer error-budgetspolicygovernancereliability -
Error Budgets advanced
A team consistently ends every window with 90% of its error budget unspent. What does that tell you?
2 min answer error-budgetsloover-investmentrisk -
Error Budgets advanced
A team consistently exhausts its error budget and continues shipping features. What has gone wrong, and what would make the budget actually function?
2 min answer crederror-budgetslogovernance -
Error Budgets advanced
How does an error budget actually change behaviour, what makes it fail in practice, and what should happen when it is exhausted?
3 min answer googlesresloerror-budget -
Failover advanced
A digital bank needs automatic failover for availability, but some operations cannot safely execute twice. Which components fail over automatically, which degrade, and which stop?
2 min answer jupiterfailoversplit-brainfinancial -
Failover advanced Multiple choice
A load balancer health check is changed from "the process answers" to "the process can reach its database and its cache". The dependency has a brief regional problem. What happens to the fleet, and why do load balancers deliberately fail open?
3 min answer health checksfail openawscorrelated failure -
Failover advanced
A regional failover completes in 90 seconds per the runbook - the database is promoted and traffic is redirected - but the application stays broken for 25 minutes. The new region's services are healthy and idle. Where is the time going?
3 min answer failoverdnsconnection poolscaching -
Failover advanced
A search platform's index-serving region becomes degraded but not fully down - elevated latency and partial errors. Should traffic fail over automatically? Analyse the risks either way.
2 min answer failovergray-failureautomationcapacity