NoSQL Stores
Key-value, document, wide-column and graph — what each buys and forbids.
5 to work through
-
intermediate
A team needs a store for derived data that is bulk-loaded from offline jobs, updated in real time, and read at very low latency. Why does this workload get its own system rather than reusing the primary online database?
2 min answer -
intermediate Multiple choice
A team proposes replacing PostgreSQL with a document store because "it scales better". Current load is 800 writes per second. What is your response?
2 min answer -
intermediate
A team wants to move a workload from PostgreSQL to a document store because "the schema keeps changing". What do you ask?
2 min answer -
advanced
A chat platform stores billions of messages in a wide-column database and hits severe tail-latency problems on hot partitions during large-community activity. Analyse why the storage choice is right but the partitioning is wrong, and what the fix looks like.
2 min answer -
advanced
A professional network computes "people you may know" from the connection graph. Some accounts have hundreds of thousands of connections. What breaks?
2 min answer
4 terms in this topic
Document Store
A store that keeps semi-structured documents — typically JSON — retrievable by key and queryable by their contents.
toolGraph Database
A store whose first-class citizens are nodes and the relationships between them, making multi-hop traversal cheap.
conceptNoSQL Stores
Storage engines that trade query flexibility and cross-entity transactions for predictable performance at scale under known access patterns.
toolWide-Column Store
A store organised as partitions of sorted rows, designed for very high write throughput and predictable single-partition reads at large scale.
Neighbouring topics
Data Architecture
General material on structuring, storing and governing data.
Relational Modelling
Normalisation, keys, constraints and the invariants a schema enforces.
Indexing
Designing indexes per query shape, and paying for them on every write.
Query Optimisation
Reading a plan, fixing statistics, and finding the real bottleneck.
Transactions & Isolation
ACID, isolation levels, and the anomalies each level permits.
Replication
Primaries, replicas, lag, and synchronous versus asynchronous durability.
Partitioning & Sharding
Splitting data across machines, and the one-way door of a partition key.
Caching Strategies
Cache-aside, read-through, write-through and where each belongs.
Cache Invalidation
Stampedes, penetration, staleness windows and versioned keys.
CQRS
Separating the write model from the read models that serve queries.
Event Sourcing
Storing the change log as the system of record, and what that costs forever.
Change Data Capture
Turning a database's replication log into a stream, and its coupling risk.
Data Warehousing
Dimensional modelling, star schemas and analytical workloads.
Data Lakes & Lakehouses
Open formats on object storage with transactional metadata on top.
ETL & ELT
Where transformation happens, and how much raw history you keep.
Streaming Data
Windowing, watermarks, late arrivals and exactly-once semantics.
Data Governance
Ownership, lineage, quality, catalogues and who may see what.
Data Lifecycle & Retention
How long data is kept, where it ages to, and how it is actually deleted.
Polyglot Persistence
Choosing a store per workload, and the operational cost of variety.