Partitioning & Sharding
Splitting data across machines, and the one-way door of a partition key.
6 to work through
-
intermediate Multiple choice
A marketplace's product database has become the bottleneck for a read-heavy workload. The team proposes sharding immediately. What lower-complexity options should be exhausted first, and in what order?
2 min answer -
advanced Multiple choice
A collaboration product is sharding its database. Options are hash of message ID, hash of user ID, or workspace ID. Which and why?
2 min answer -
advanced
A messaging platform stores messages partitioned by channel identifier. Some channels are enormously more active than others. Fix the partitioning.
2 min answer -
advanced
A multi-tenant SaaS platform runs thousands of tenants on a shared relational database, but a handful of tenants now hold most of the data. Should it use shared tables, tenant partitioning, dedicated databases, sharding or a hybrid?
2 min answer -
advanced
A workspace product stores every document as a tree of small "block" rows in one enormous table. Growth makes the table unmanageable. Design the sharding strategy - and identify the choice that is effectively irreversible.
2 min answer -
advanced
Your geospatial data is partitioned by city. One city generates 40% of all traffic. What are your options?
2 min answer
10 terms in this topic
Consistent Hashing
A partitioning scheme where adding or removing a node moves only a small fraction of keys, instead of remapping everything.
case-studyDiscord: Hot Partitions at Trillions of Messages
Discord partitions messages by channel and time bucket, because a single very busy channel would otherwise concentrate load on one partition.
conceptHot Key
A single key or narrow key range receiving a disproportionate share of traffic, so one partition saturates while the rest of the cluster is idle - th…
conceptHot Key
A single key or partition receiving a disproportionate share of traffic, so that a well-balanced key space still produces one overloaded node.
conceptHot Partition
One partition receiving disproportionate traffic, so the system saturates at a fraction of its aggregate capacity.
conceptPartitioning and Sharding
Splitting data by a key — within one database for manageability, or across databases for capacity and isolation.
practiceShard Key
The attribute deciding which partition a row belongs to - the single most consequential and least reversible choice in a partitioned data architecture.
patternSharding in Practice
Splitting data across independent stores, how to choose the key, and why resharding is the operation nobody plans for.
patternTenant Placement
Routing each tenant to a shared pool or a dedicated database according to its size and requirements, with an online migration path between them - the…
conceptWrite Amplification
One logical write producing many physical writes - through fan-out, indexes, replication or storage-engine mechanics - and why it decides scaling limits.
Neighbouring topics
Data Architecture
General material on structuring, storing and governing data.
Relational Modelling
Normalisation, keys, constraints and the invariants a schema enforces.
NoSQL Stores
Key-value, document, wide-column and graph — what each buys and forbids.
Indexing
Designing indexes per query shape, and paying for them on every write.
Query Optimisation
Reading a plan, fixing statistics, and finding the real bottleneck.
Transactions & Isolation
ACID, isolation levels, and the anomalies each level permits.
Replication
Primaries, replicas, lag, and synchronous versus asynchronous durability.
Caching Strategies
Cache-aside, read-through, write-through and where each belongs.
Cache Invalidation
Stampedes, penetration, staleness windows and versioned keys.
CQRS
Separating the write model from the read models that serve queries.
Event Sourcing
Storing the change log as the system of record, and what that costs forever.
Change Data Capture
Turning a database's replication log into a stream, and its coupling risk.
Data Warehousing
Dimensional modelling, star schemas and analytical workloads.
Data Lakes & Lakehouses
Open formats on object storage with transactional metadata on top.
ETL & ELT
Where transformation happens, and how much raw history you keep.
Streaming Data
Windowing, watermarks, late arrivals and exactly-once semantics.
Data Governance
Ownership, lineage, quality, catalogues and who may see what.
Data Lifecycle & Retention
How long data is kept, where it ages to, and how it is actually deleted.
Polyglot Persistence
Choosing a store per workload, and the operational cost of variety.