Distributed Systems

General material on partial failure, coordination and distributed reasoning.

11Questions
27Flashcards
19Terms
Questions

11 to work through

  1. intermediate

    An order service must notify inventory, billing, shipping and analytics when an order is placed. Synchronous calls or events? Justify your choice per consumer.

    2 min answer
  2. intermediate

    Design an order submission API that is safe when the client cannot tell whether its request succeeded. What exactly do you store, and when?

    2 min answer
  3. intermediate Multiple choice

    You move a user profile service to eventual consistency and support tickets start arriving: users update their name and the old one is still shown. Fix it without abandoning the architecture.

    2 min answer
  4. advanced

    A 43-second network partition caused GitHub over 24 hours of degraded service in 2018. How does a 43-second event become a day-long incident?

    2 min answer
  5. advanced Multiple choice

    A card payment authorisation service runs active-active across two regions. A network partition splits them. Do you keep accepting authorisations, and what breaks either way?

    2 min answer
  6. advanced

    A downstream service slows from 50 ms to 3 s. Within two minutes every service in the request path is down, including ones that do not call it. Explain the mechanism and how you would have prevented it.

    2 min answer
  7. advanced

    A payments platform sees a 20x increase in transaction attempts during a major commerce event. Idempotency, rate limiting, queueing, fraud checks, database contention and a downstream provider all interact. What is the correct ordering of these controls on the request path, and why does ordering matter more than any single control?

    2 min answer
  8. advanced

    Design the connection layer for a chat platform holding tens of millions of concurrent WebSocket connections, where users belong to communities ranging from three people to a million. What are the main architectural decisions, and which one causes the most incidents?

    2 min answer
  9. advanced

    Design the system that matches a rider request to a nearby driver. What are the hard parts?

    2 min answer
  10. advanced

    DynamoDB has evolved through adaptive capacity, on-demand mode and global tables. Which failure mode did each address, and what does the sequence teach about designing partitioned systems?

    2 min answer
  11. advanced

    Showing which users are currently online seems trivial and is one of the most expensive features in a chat product. Explain.

    2 min answer
Terminology

19 terms in this topic

concept

Backpressure

A mechanism by which a component under load tells its callers to slow down, rather than accepting work it cannot complete.

concept

CAP Theorem

During a network partition a distributed system must choose between consistency and availability; it cannot have both.

pattern

Circuit Breaker

A proxy that stops calling a failing dependency after a failure threshold, failing fast instead, and periodically tests whether it has recovered.

tool

Event Stream

An append-only, retained log of events that many independent consumers read at their own position, and can re-read.

concept

Eventual Consistency

A guarantee that replicas will converge to the same value if updates stop, with no bound on how long reads may be stale.

pattern

Exponential Backoff

Increasing the wait between retries geometrically, with random jitter, so that failures do not synchronise into a stampede.

concept

Fan-Out

One incoming request causing many outgoing ones, which multiplies both load and tail latency.

concept

Fault Tolerance

Continuing to operate correctly despite the failure of some components, by design rather than by luck.

practice

Graceful Degradation

Continuing to deliver reduced but useful function when a dependency fails, instead of failing the whole request.

concept

Idempotency

The property that performing an operation many times has the same effect as performing it once.

pattern

Leader Election

The process by which a group of nodes agrees which one of them is currently in charge of a task that must not run twice.

tool

Message Queue

A store that holds messages until a consumer processes them, decoupling producer availability and rate from consumer availability and rate.

concept

Reconnect Storm

The synchronised reconnection of a large population of clients after a disruption, which routinely causes a larger outage than the disruption itself.

pattern

Saga

A sequence of local transactions across services where each step has a compensating action that semantically undoes it if a later step fails.

concept

Scalability

The ability to handle growing load by adding resources, ideally with cost rising no faster than the load.

concept

Service Discovery

The mechanism by which a caller finds a currently healthy network address for a service whose instances are ephemeral.

concept

Thundering Herd

A large number of clients acting simultaneously because they were synchronised by a shared event, producing a spike that the steady-state design neve…

pattern

Timeout Budget

Assigning a request an overall deadline at the edge and passing the remaining time down each hop, so no service works on something already out of time.

protocol

Two-Phase Commit

A blocking protocol for atomic commit across several resources: a coordinator asks all participants to prepare, then tells them all to commit or abort.

Distributed Systems

Neighbouring topics

CAP & PACELC

What you must give up during a partition, and the latency choice the rest of the time.

4 quiz 13 cards 3 terms

Consistency Models

Linearizable, sequential, causal, eventual, and the session guarantees between them.

5 quiz 10 cards 7 terms

Idempotency

Making an operation safe to repeat, because a client that times out cannot know.

5 quiz 14 cards 3 terms

Retries & Backoff

Exponential backoff, jitter, retry budgets, and how retries become the outage.

4 quiz 7 cards 2 terms

Timeouts & Deadlines

Per-hop timeouts that do not compose, and the deadline budget that replaces them.

5 quiz 9 cards 6 terms

Circuit Breakers

Failing fast on a broken dependency, and what you fail fast to.

6 quiz 9 cards 5 terms

Backpressure & Flow Control

Telling callers to slow down instead of buffering into congestion collapse.

4 quiz 10 cards 5 terms

Load Shedding

Rejecting some work deliberately so the rest can be served correctly.

6 quiz 10 cards 6 terms

Bulkheads & Isolation

Partitioning resources so one dependency cannot starve the others.

5 quiz 10 cards 6 terms

Leader Election

Agreeing who is in charge, and fencing the one who no longer is.

6 quiz 11 cards 5 terms

Consensus Protocols

Raft, Paxos and quorums — what they guarantee and what they cost.

2 quiz 7 cards 2 terms

Distributed Locking

Mutual exclusion across machines, and why it is harder than it looks.

4 quiz 9 cards 6 terms

Distributed Transactions

Two-phase commit, its blocking failure mode, and when it is still reasonable.

4 quiz 15 cards 5 terms

Sagas & Compensation

Replacing atomicity with semantic undo, and ordering the irreversible steps last.

6 quiz 11 cards 5 terms

Service Discovery

Finding a healthy address for something whose instances are ephemeral.

4 quiz 15 cards 4 terms

Messaging & Queues

Decoupling producer from consumer, and the semantics that come with it.

4 quiz 13 cards 3 terms

Event Streaming

Retained ordered logs, consumer offsets, partitions and replay.

6 quiz 16 cards 4 terms

Clocks & Ordering

Why wall clocks lie, and how logical clocks and versions restore order.

4 quiz 13 cards 2 terms

Failure Modes

Slow rather than down, partial, grey, and failing while reporting success.

6 quiz 13 cards 9 terms