Quiz
2777 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2777
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture87
Reliability & Resilience99
Observability92
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture100
Legacy Modernization92
AI-Era Architecture96
Software Architecture & Engineering93
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture102
Data Platform Architecture98
Streaming & Real-Time Data93
Data Governance & Semantics91
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture99
Assurance, Audit & Model Risk98
2777 questions.
-
DR Testing intermediate
An enterprise runs an annual disaster recovery test that always succeeds, yet the team has low confidence in real recovery. What is likely wrong with the test?
2 min answer dr-testingdrillsrealismenterprise -
DR Testing advanced
In January 2017 GitLab lost roughly six hours of database data after an engineer deleted a directory on the wrong host during replication troubleshooting, and then found that several backup and replication mechanisms had silently not been working. What signal would have revealed the broken backups beforehand, and why did nothing report them?
3 min answer gitlabbackupsrestore testingsilent failure -
DR Testing advanced
Your DR plan is pilot light with a 30-minute RTO. What would you test, and what will the test probably reveal?
2 min answer drtestingcontrol-plane -
DR Testing advanced
Zoom's daily meeting participants rose from roughly 10 million in December 2019 to about 300 million by April 2020. Your platform has a plausible 10x surge ahead and a disaster-recovery plan last tested by tabletop review two years ago. Sequence the move from tabletop to a genuinely tested recovery capability, without an outage.
4 min answer zoomdisaster recoverydr testingsurge -
Distributed Locking advanced
A media platform uses a distributed lock to stop two workers rendering the same expensive export. Workers sometimes crash while holding the lock, and occasionally two workers process the same job anyway. Diagnose both problems and redesign.
2 min answer distributed-lockingleasesfencingidempotency -
Distributed Locking advanced
A team proposes a distributed lock to stop two workers processing the same transaction. Workers sometimes crash or pause while holding the lock. What is the safer design, and when is a lock genuinely necessary?
2 min answer jupiterdistributed-lockfencingidempotency -
Distributed Locking advanced Multiple choice
A team uses a distributed lock in a key-value store to ensure only one worker processes a job. Occasionally two workers process the same job. Explain.
2 min answer lockingcorrectnessfencing -
Distributed Locking advanced
Three designs need distributed locks: a nightly report, a per-customer state machine, and a global config reload. For each, is a lock the right answer?
2 min answer lockingpartitioningidempotencydesign -
Distributed Systems advanced
A 43-second network partition caused GitHub over 24 hours of degraded service in 2018. How does a 43-second event become a day-long incident?
2 min answer failoversplit-brainconsistencycase-study -
Distributed Systems advanced Multiple choice
A card payment authorisation service runs active-active across two regions. A network partition splits them. Do you keep accepting authorisations, and what breaks either way?
2 min answer capconsistencypaymentsavailability -
Distributed Systems advanced
A downstream service slows from 50 ms to 3 s. Within two minutes every service in the request path is down, including ones that do not call it. Explain the mechanism and how you would have prevented it.
2 min answer cascading-failureretriestimeoutsresilience -
Distributed Systems advanced
A payments platform sees a 20x increase in transaction attempts during a major commerce event. Idempotency, rate limiting, queueing, fraud checks, database contention and a downstream provider all interact. What is the correct ordering of these controls on the request path, and why does ordering matter more than any single control?
2 min answer razorpaypaymentssurgeidempotency