Term Kind Topic What it is
Multi-Leader Replication Multi-Master, Active-Active Replication pattern Replication Accepting writes at more than one node and replicating between them, which removes the single-primary bottleneck and introduces write conflicts.
Netflix's Recommendation Architecture case-study Data Architecture Netflix splits personalisation into offline, nearline and online layers so that expensive computation happens ahead of time and the request path stays fast.
Normalisation Normal Forms practice Relational Modelling Organising a schema so each fact is stored exactly once, removing the update anomalies that duplication creates.
NoSQL Stores Non-Relational Databases concept NoSQL Stores Storage engines that trade query flexibility and cross-entity transactions for predictable performance at scale under known access patterns.
OLTP vs OLAP concept Data Warehousing Two workload shapes with opposite requirements — many small indexed transactions versus few large scans and aggregations — which is why they belong in different stores.
Operational vs Analytical Store concept Polyglot Persistence The separation between the store serving the application's transactions and the one serving reporting and analysis, and the mechanism connecting them.
Partial Index Filtered Index concept Indexing An index built over only the rows matching a predicate, so it is far smaller and cheaper to maintain than a full index.
Partitioning and Sharding concept Partitioning & Sharding Splitting data by a key — within one database for manageability, or across databases for capacity and isolation.
Pattern Scope Discipline Apply the Pattern to a Use Case, Not the System concept CQRS The rule that heavyweight data patterns - CQRS, event sourcing, materialised read models - belong to the specific part of a domain that needs them, never to a whole product.
Pinterest's MySQL Sharding case-study Data Architecture Pinterest sharded MySQL by embedding the shard ID inside every primary key, making any object's location computable from its ID alone with no lookup service.
Pipeline Orchestration practice ETL & ELT Coordinating the execution of data tasks by dependency rather than by clock, with retries, backfill and observability built in.
Polyglot Persistence concept Polyglot Persistence Using different storage engines for different workloads within one system — justified by genuine workload divergence, and frequently not.
Projection pattern CQRS The process that consumes changes from the write side and maintains a read model, and the component where most CQRS bugs live.
Query Optimisation practice Query Optimisation Making the database do less work — usually by removing round trips and rows rather than by rewriting clever SQL.
Query Plan Execution Plan, EXPLAIN concept Data Architecture The database's chosen strategy for executing a query, and the first thing to look at when one is slow.
Query-Based CDC Polling CDC, Timestamp-Based CDC pattern Change Data Capture Detecting changes by repeatedly querying for rows modified since the last run — simple, universally available, and lossy in specific ways.
Read Model Query Model, Projection Store pattern CQRS A data structure shaped for a specific query rather than for the domain, maintained separately from the write model.
Read Replica pattern Data Architecture A copy of a database that receives changes from the primary and serves read-only queries, spreading read load.
Relational Modelling practice Relational Modelling Designing a schema around entities, relationships and enforced constraints — still the correct default for most transactional systems.
Replica Lag Routing Read Routing, Primary Pinning practice Replication Deciding, per read, whether it may be served by an asynchronous replica - the operational discipline that makes read scaling safe.
Replication concept Replication Keeping copies of data on multiple machines for availability, read scaling and locality — and the lag that comes with all three.
Replication in Practice Read Replicas, Multi-Primary pattern Replication Copying data across nodes for availability and read capacity, and the lag that turns into user-visible bugs if session guarantees are not designed in.
Replication Lag metric Replication How far behind a replica is, measured in time or in log position — the quantity that determines how stale a replica read can be.
Request Coalescing Request Collapsing, Single-Flight practice Caching Strategies Collapsing many concurrent requests for the same missing resource into a single upstream fetch, with the rest waiting on its result.
Reservation Ledger Stock Ledger, Reserve-Then-Commit pattern Transactions & Isolation Modelling stock as an append-only sequence of reservations, releases and allocations rather than as a mutable count - so the number has a history, can be reconciled, and does not become a single point of contention.
Salesforce's Metadata-Driven Multi-Tenancy case-study Data Architecture Salesforce serves every customer from shared infrastructure with a single physical schema, storing customer-specific data structures as metadata rather than as separate tables.
Schema Evolution concept Data Lakes & Lakehouses Changing a table's structure over time while keeping existing data readable and existing consumers working.
Shard Key Partition Key practice Partitioning & Sharding The attribute deciding which partition a row belongs to - the single most consequential and least reversible choice in a partitioned data architecture.
Sharding Horizontal Partitioning pattern Data Architecture Splitting one dataset across multiple independent databases by a partition key, so that each holds a disjoint subset.
Sharding in Practice Horizontal Partitioning pattern Partitioning & Sharding Splitting data across independent stores, how to choose the key, and why resharding is the operation nobody plans for.
Slack Flannel: Caching at the Edge for a Chat Client Flannel case-study Caching Strategies Slack pushed user and channel metadata into an application-aware edge cache because clients were downloading enormous amounts of it on every connection.
Snapshotting pattern Event Sourcing Periodically storing an aggregate's computed state so it can be loaded without replaying its entire event history.
Star Schema pattern Data Warehousing A dimensional model with one central fact table of measurements surrounded by denormalised dimension tables describing them.
Streaming Data Architecture concept Streaming Data Processing unbounded data continuously, where the central problems are time semantics, lateness and state rather than throughput.
Surrogate Key Synthetic Key concept Relational Modelling A system-generated identifier with no business meaning, used as the primary key instead of a naturally occurring business value.
Synchronous vs Asynchronous Replication concept Replication Whether a write is acknowledged only after a replica has it, trading write latency against the amount of data a failure can lose.
Tenant Placement Hybrid Tenancy, Shared-to-Dedicated Migration pattern Partitioning & Sharding Routing each tenant to a shared pool or a dedicated database according to its size and requirements, with an online migration path between them - the hybrid model that neither all-shared nor all-dedicated can match.
Time Travel Data Versioning, Snapshot Query concept Data Lakes & Lakehouses Querying a table as it existed at a previous version or timestamp, made possible by keeping the metadata and files of prior commits.
Transactions and Isolation ACID Isolation Levels concept Transactions & Isolation What a database guarantees when concurrent transactions touch the same data — and the anomalies each level permits.
Two-Phase Locking 2PL protocol Transactions & Isolation The concurrency control protocol behind serializable isolation — acquire locks in a growing phase, release only in a shrinking phase, never interleaving the two.
Uber H3: Hexagonal Spatial Indexing H3, Hexagonal Hierarchical Index case-study Indexing Uber built and open-sourced a hexagonal grid index because uniform neighbour distance matters when you are analysing supply and demand across space.
Uber's H3 Spatial Index case-study Data Architecture Uber indexes the world with hexagons rather than squares, because uniform neighbour distance makes supply, demand and pricing computations correct as well as fast.
Versioned Dataset Swap Atomic Pointer Flip, Generation Swap, Blue-Green Data pattern Data Architecture Publishing a regenerated dataset as a complete new version alongside the live one and flipping a serving pointer atomically, so readers never observe a mixture of generations and rollback is a pointer flip rat…
Wide-Column Store tool NoSQL Stores A store organised as partitions of sorted rows, designed for very high write throughput and predictable single-partition reads at large scale.
Windowing concept Streaming Data Grouping an unbounded stream into finite chunks so aggregation can produce results, defined over event time rather than arrival time.
Workload Class Isolation Compute Separation by Class, Interactive vs Batch Pools pattern Data Lakes & Lakehouses Running interactive, scheduled and batch analytical workloads on separate compute over shared storage, so that a long-running job cannot make an analyst's query unpredictable.
Write Amplification concept Partitioning & Sharding One logical write producing many physical writes - through fan-out, indexes, replication or storage-engine mechanics - and why it decides scaling limits.
Write Skew concept Transactions & Isolation An anomaly where two transactions each read a set of rows, make disjoint writes based on what they read, and together violate an invariant neither could have broken alone.
Write-Ahead Log WAL, Commit Log concept Data Architecture Recording every change to a durable sequential log before applying it, so that a crash can be recovered by replaying the log.
Write-Amplification Audit Index Cost Review, Write Path Accounting practice Query Optimisation Accounting for everything a single logical write actually costs - indexes, replication, triggers, deletion of old rows - before concluding that a database needs more capacity.