Data Lifecycle & Retention
How long data is kept, where it ages to, and how it is actually deleted.
4 to work through
-
advanced
A GDPR erasure request arrives for a customer. Where does their data actually live, and what makes this expensive to retrofit?
2 min answer -
advanced
A consumer fintech accumulates transaction, behavioural and derived data. How should retention, archival, deletion and regulatory holds be designed so they do not conflict?
2 min answer -
advanced
A customer exercises their right to erasure. List every place their data plausibly exists in a mature architecture and how each is handled.
3 min answer -
advanced
A file storage platform must permanently delete a user's data on request, while maintaining backups, replicas, caches and search indexes. What makes this genuinely difficult, and how is it designed?
3 min answer
5 terms in this topic
Crypto-Shredding
Encrypting each data subject's personal data under a key unique to them, so that a deletion request is satisfied by destroying the key rather than by…
practiceCrypto-Shredding
Making data permanently unreadable by destroying its encryption key, so immutable copies in backups and archives are erased without being modified.
practiceData Lifecycle
Managing data from creation through tiering to deletion — the discipline that keeps storage cost and legal exposure bounded.
conceptHot, Warm and Cold Data
Classifying data by how frequently and how urgently it is accessed, so each tier can be stored on media priced for that access pattern.
patternIdentity-Record Separation
Holding regulated business records under a pseudonymous reference and the reference-to-person mapping in a separate governed store, so a retention ob…
Neighbouring topics
Data Architecture
General material on structuring, storing and governing data.
Relational Modelling
Normalisation, keys, constraints and the invariants a schema enforces.
NoSQL Stores
Key-value, document, wide-column and graph — what each buys and forbids.
Indexing
Designing indexes per query shape, and paying for them on every write.
Query Optimisation
Reading a plan, fixing statistics, and finding the real bottleneck.
Transactions & Isolation
ACID, isolation levels, and the anomalies each level permits.
Replication
Primaries, replicas, lag, and synchronous versus asynchronous durability.
Partitioning & Sharding
Splitting data across machines, and the one-way door of a partition key.
Caching Strategies
Cache-aside, read-through, write-through and where each belongs.
Cache Invalidation
Stampedes, penetration, staleness windows and versioned keys.
CQRS
Separating the write model from the read models that serve queries.
Event Sourcing
Storing the change log as the system of record, and what that costs forever.
Change Data Capture
Turning a database's replication log into a stream, and its coupling risk.
Data Warehousing
Dimensional modelling, star schemas and analytical workloads.
Data Lakes & Lakehouses
Open formats on object storage with transactional metadata on top.
ETL & ELT
Where transformation happens, and how much raw history you keep.
Streaming Data
Windowing, watermarks, late arrivals and exactly-once semantics.
Data Governance
Ownership, lineage, quality, catalogues and who may see what.
Polyglot Persistence
Choosing a store per workload, and the operational cost of variety.