1. Consistency Models advanced

    A food-delivery platform must match orders to nearby couriers while supply and demand change every few seconds. Which information needs strong consistency, which can be stale, and how should dispatch handle a courier who accepted a job while their availability record was out of date?

    2 min answer swiggydispatchgeospatialstaleness
  2. Consistency Models advanced Multiple choice

    For a shopping cart at very large scale, would you choose a strongly consistent store or an always-writeable one with conflict resolution?

    2 min answer amazondynamoavailabilitytradeoffs
  3. Consistency Models advanced

    Two people are editing the same page in a collaborative workspace. One loses connectivity for several minutes while continuing to type, then reconnects. What should happen, and what consistency guarantee is actually required?

    2 min answer consistencycollaborationofflinenotion
  4. Distributed Locking advanced

    A media platform uses a distributed lock to stop two workers rendering the same expensive export. Workers sometimes crash while holding the lock, and occasionally two workers process the same job anyway. Diagnose both problems and redesign.

    2 min answer distributed-lockingleasesfencingidempotency
  5. Distributed Locking advanced

    A team proposes a distributed lock to stop two workers processing the same transaction. Workers sometimes crash or pause while holding the lock. What is the safer design, and when is a lock genuinely necessary?

    2 min answer jupiterdistributed-lockfencingidempotency
  6. Distributed Locking advanced Multiple choice

    A team uses a distributed lock in a key-value store to ensure only one worker processes a job. Occasionally two workers process the same job. Explain.

    2 min answer lockingcorrectnessfencing
  7. Distributed Locking advanced

    Three designs need distributed locks: a nightly report, a per-customer state machine, and a global config reload. For each, is a lock the right answer?

    2 min answer lockingpartitioningidempotencydesign
  8. Distributed Systems advanced

    A 43-second network partition caused GitHub over 24 hours of degraded service in 2018. How does a 43-second event become a day-long incident?

    2 min answer failoversplit-brainconsistencycase-study
  9. Distributed Systems advanced Multiple choice

    A card payment authorisation service runs active-active across two regions. A network partition splits them. Do you keep accepting authorisations, and what breaks either way?

    2 min answer capconsistencypaymentsavailability
  10. Distributed Systems advanced

    A downstream service slows from 50 ms to 3 s. Within two minutes every service in the request path is down, including ones that do not call it. Explain the mechanism and how you would have prevented it.

    2 min answer cascading-failureretriestimeoutsresilience
  11. Distributed Systems advanced

    A payments platform sees a 20x increase in transaction attempts during a major commerce event. Idempotency, rate limiting, queueing, fraud checks, database contention and a downstream provider all interact. What is the correct ordering of these controls on the request path, and why does ordering matter more than any single control?

    2 min answer razorpaypaymentssurgeidempotency
  12. Distributed Systems advanced

    Design the connection layer for a chat platform holding tens of millions of concurrent WebSocket connections, where users belong to communities ranging from three people to a million. What are the main architectural decisions, and which one causes the most incidents?

    2 min answer websocketsgatewayshardingpresence