Polyglot Persistence
Choosing a store per workload, and the operational cost of variety.
6 to work through
-
intermediate
A developer platform accumulates a relational database, an object store, a search index, a cache, a queue and a time-series store. When is this justified polyglot persistence, and when is it accidental sprawl?
2 min answer -
intermediate
A team wants to add a graph database for a "people you may know" feature. The system currently runs on one PostgreSQL instance. Evaluate.
2 min answer -
advanced
A delivery platform migrated parts of its data estate to CockroachDB and other parts to Aurora. What evidence distinguishes workloads that justify a distributed SQL database from those better served by a single-writer managed Postgres?
2 min answer -
advanced
A trading platform is considering a relational database, a key-value store, a time-series store and a search engine. What justifies each, and what is the cost of running four?
2 min answer -
advanced
Four services want four different databases: Postgres, MongoDB, Cassandra and Neo4j. What do you say?
2 min answer -
advanced
Your search index and your database disagree — some products appear in search that were deleted, and some new ones never appear. How do you make this reliable?
2 min answer
4 terms in this topic
Database per Service
Each service owning its own datastore, with no other service reading or writing it directly.
conceptDerived Store Discipline
The rule that every fact has exactly one authoritative store and every other copy is explicitly derived, rebuildable and populated by a documented pi…
conceptOperational vs Analytical Store
The separation between the store serving the application's transactions and the one serving reporting and analysis, and the mechanism connecting them.
conceptPolyglot Persistence
Using different storage engines for different workloads within one system — justified by genuine workload divergence, and frequently not.
Neighbouring topics
Data Architecture
General material on structuring, storing and governing data.
Relational Modelling
Normalisation, keys, constraints and the invariants a schema enforces.
NoSQL Stores
Key-value, document, wide-column and graph — what each buys and forbids.
Indexing
Designing indexes per query shape, and paying for them on every write.
Query Optimisation
Reading a plan, fixing statistics, and finding the real bottleneck.
Transactions & Isolation
ACID, isolation levels, and the anomalies each level permits.
Replication
Primaries, replicas, lag, and synchronous versus asynchronous durability.
Partitioning & Sharding
Splitting data across machines, and the one-way door of a partition key.
Caching Strategies
Cache-aside, read-through, write-through and where each belongs.
Cache Invalidation
Stampedes, penetration, staleness windows and versioned keys.
CQRS
Separating the write model from the read models that serve queries.
Event Sourcing
Storing the change log as the system of record, and what that costs forever.
Change Data Capture
Turning a database's replication log into a stream, and its coupling risk.
Data Warehousing
Dimensional modelling, star schemas and analytical workloads.
Data Lakes & Lakehouses
Open formats on object storage with transactional metadata on top.
ETL & ELT
Where transformation happens, and how much raw history you keep.
Streaming Data
Windowing, watermarks, late arrivals and exactly-once semantics.
Data Governance
Ownership, lineage, quality, catalogues and who may see what.
Data Lifecycle & Retention
How long data is kept, where it ages to, and how it is actually deleted.