At 21:40 a schema change on the legacy order database makes the change-data-capture connector feeding the new analytics store crash-loop. At 23:10 the connector is still down and nobody has been paged, because the analytics store is not customer-facing. At 04:55 the order database refuses writes — its data volume is full. What failed, and which design decision made it possible?
Show the full answer Hide the answer
The trigger
The schema change is the trigger, not the cause. A column type that the connector's converter cannot map makes it fail on the first row it reads after the change, restart, read the same row and fail again. From the database's point of view nothing is wrong: the subscriber has simply not confirmed any progress since 21:40.
Why it propagated
A logical replication slot is a promise. The database retains every write-ahead log segment from the slot's restart position forward until the consumer confirms it has them, because that is the only way the consumer can resume without losing changes. A stalled consumer therefore turns the primary's log directory into an unbounded buffer for the consumer's downtime.
The arithmetic is unforgiving. A busy order database generating 15 MB of WAL a second on average retains about 54 GB an hour. From 21:40 to 04:55 is 7 hours 15 minutes, which is roughly 390 GB — enough to consume the headroom on a volume sized for a few days of normal churn. When the volume fills, Postgres cannot write, so it stops accepting writes. The customer-facing database is down because a non customer-facing consumer stopped reading.
The design decision that made it possible is the default. max_slot_wal_keep_size has defaulted to -1,
meaning unlimited retention, since the setting was introduced in PostgreSQL 13. Leaving it there means the
database is configured to protect the consumer's resumability at the cost of its own availability, which is
exactly backwards for an offload whose entire purpose was to take load off the legacy system.
Why detection lagged
Because the alerting followed the data's importance rather than the dependency's direction. The analytics store is not customer-facing, so its lag was a business-hours ticket. The signal that mattered was not "the analytics store is stale" but "the primary is retaining WAL", and nobody owned that metric — the platform team watched disk free space as a percentage with a 24-hour trend, which is useless against a 54 GB-an-hour slope, and the data team watched freshness without knowing what staleness cost the primary.
The structural fix versus the tempting local fix
The tempting fix is a bigger volume. It converts a 7-hour fuse into a 30-hour fuse and leaves the coupling intact, so the same incident recurs on a long weekend.
The structural fixes, in order:
- Set
max_slot_wal_keep_sizeto a finite value sized at the longest connector outage you are willing to absorb, say 4 hours of peak WAL. Past that the database invalidates the slot and keeps serving. The consumer then needs reseeding from a snapshot, which is a day of work for the data team and is categorically better than a write outage. - Alert on retained WAL expressed in hours of headroom, not bytes and not percentage free, with the page going to whoever owns the primary.
- Make connector schema compatibility part of the database's change process, so a column type change is tested against the connector before it ships.
Common weak answers
- "Monitor the connector more closely." Names no signal and no threshold. The connector was visibly broken for seven hours and it did not matter, because the alert that was missing was on the database.
- "Use a queue instead of CDC." Moves the retention problem to the queue and adds a dual-write correctness problem. The coupling is retention, not the transport.
- "The schema change should have been reviewed." True and insufficient. Any connector outage of a few hours produces this, including a deploy or a credential expiry.
The general lesson
An offload added to protect a legacy system becomes a write dependency of it the moment the offload's progress controls the legacy system's log retention. Every CDC consumer is a liveness requirement on its source unless retention is explicitly bounded, and the bound is the one setting most teams never change.