1. Failure Modes advanced

    What happens when a globally distributed real-time communication service loses an entire region - and how should DNS, health checks, traffic steering, failover, connection draining and data consistency interact?

    3 min answer failure-modesregional-failurefailoverdns
  2. Idempotency advanced Multiple choice

    A billing platform sends a webhook, the customer's endpoint times out, and the customer's system retries the operation several times. How should idempotency keys, durable event state, delivery attempts, ordering and reconciliation prevent duplicate billing?

    2 min answer chargebeewebhooksidempotencyat-least-once
  3. Idempotency intermediate Multiple choice

    A marketplace checkout calls a payment API. Mobile clients on poor networks time out and retry, and some customers are charged twice. Where exactly must the deduplication record be written for this to be fixed?

    2 min answer idempotencyretriespaymentscorrectness
  4. Idempotency advanced

    A payments API receives the same charge request twice because a mobile client retried after a timeout. How should idempotency keys be scoped, stored, expired and validated, and what happens when the first attempt is still in flight?

    2 min answer stripeidempotencyretriespayments
  5. Idempotency advanced

    An API platform sends messages on behalf of external developers whose clients retry aggressively after network timeouts. Design idempotency so that a retried request never sends a second message - and state precisely where the guarantee begins and ends.

    2 min answer idempotencyretriesapi-designtwilio
  6. Idempotency advanced Multiple choice

    Mobile clients on poor networks retry payment requests after timeouts. Where must idempotency be enforced, what must be durable before responding, and what is the most common scoping mistake?

    2 min answer phonepeidempotencyretriesmobile
  7. Leader Election advanced

    A leader-based coordination service loses contact with its followers while the leader itself remains healthy and continues serving. What failure modes emerge, and what prevents split-brain?

    2 min answer polygonleader-electionsplit-brainquorum
  8. Leader Election advanced

    A leader-based database cluster loses network connectivity between the leader and its followers, but the leader is still running and still accepting writes from application servers that can reach it. What happens, and how does correct leader election prevent it?

    3 min answer leader-electionsplit-brainquorumfencing
  9. Leader Election advanced

    A nightly job occasionally runs twice, producing duplicate charges. The team proposes a distributed lock. What do you say?

    2 min answer leader-electionlockingidempotencyfencing
  10. Leader Election advanced

    A team designs active-active across two regions with automatic failover based on health checks. What is wrong?

    2 min answer split-brainquorummulti-region
  11. Leader Election advanced

    Your cluster fails over spuriously under load, but a real leader failure takes 45 seconds to detect. How do you resolve the tension?

    2 min answer leader-electionfailoverdetectiontuning
  12. Leader Election intermediate Multiple choice

    Your company runs two data centres and wants automatic database failover between them. What is missing?

    2 min answer quorumsplit-brainmulti-regiondesign