An MQTT 5 broker cluster serving 120,000 sensors starts alerting on disk at 03:00 every night, and the oldest of those alerts is three weeks old. Publish rate is flat at roughly 900 messages per second, connected clients sit at 118,000, and the broker reports 340,000 sessions and rising. No messages are being dropped yet. What failed, and which design decision made it possible?
Show the full answer Hide the answer
The trigger
A firmware release six weeks ago changed the client identifier from the device serial number to the serial plus a random suffix. The change was made to work around a reconnect race: the broker refused a new connection while the previous session was still being torn down, and a fresh identifier made the error go away.
In MQTT a session is keyed by the client identifier, so a new identifier on every boot creates a new session and orphans the previous one. These devices connect with Clean Start = 0 and a Session Expiry Interval of seven days, which is exactly what makes downlink work for a sensor that sleeps between reports.
Why it propagates instead of settling
An orphaned session is not dormant storage. It keeps the device's subscriptions, so the broker goes on queueing every QoS 1 message matching them for a subscriber that will never reconnect, until the expiry interval elapses. The fleet reboots on a weekly maintenance window, so each device contributes a new orphan faster than old ones expire.
The arithmetic is the whole story. Roughly 220,000 orphan sessions, each accumulating a few hundred queued commands and acknowledgement state at a few hundred bytes, is 220,000 × 400 × 300 B, or about 26 GB of messages with no recipient. The nightly alert is the compaction job failing to keep up, not a storage sizing error.
Why detection lagged
The dashboards chart connected clients and publish throughput, and both are healthy. Session count is a broker gauge almost nobody plots, because in a normal fleet it tracks the device count so closely that it looks redundant. The signal that names this in one glance is sessions minus connections, which went from near zero to 220,000 over six weeks.
The structural fix, and the tempting local one
The tempting fix is to cut the Session Expiry Interval to an hour. Disk recovers that night, and every command issued while a device sleeps is now discarded, which is the feature the persistent session was bought for.
What actually works:
- Derive the client identifier from hardware identity and never vary it. MQTT 5 already defines takeover: a new connection with the same identifier displaces the old one and continues the session. The reconnect race was a client bug, not a protocol gap.
- Set Session Expiry Interval from the measured p99 offline window, not a round number someone liked.
- Cap the queue per session and set a Message Expiry Interval on every downlink publish, so an instruction that is no longer useful dies instead of waiting.
- Alert on sessions minus connections above 5% of the fleet.
Common weak answers
- "Send downlink at QoS 0." The queue disappears and so does delivery to any sleeping device, which converts a storage problem into a silent control-plane problem.
- "Add brokers." Cost then grows with the leak. Capacity planning against an unbounded identifier space never converges.
- "Set Clean Start = 1 on every connect." Disk recovers and the fleet quietly loses every command sent while a device was away, with no error anywhere.
The general lesson
Any identifier a client picks becomes a primary key in server-side storage, and a client that can pick a new one can allocate storage without limit. The same shape appears with consumer-group names, idempotency keys scoped per boot, and per-session temporary topics.