Store-and-Forward Queue
also called On-Device Send Queue, Deferred Send Buffer
A durable queue on the device that accepts writes while the network is absent and drains them in order when it returns - with an explicit policy for what happens when it fills or an entry is rejected.
A courier marks a parcel delivered in a basement with no signal. The parcel is delivered whether or not a server knows, so the application has two options: refuse the action, or accept it locally and transmit later. Every usable field application chooses the second, and the queue that holds those accepted-but-unsent actions is where the design either survives a bad week or loses a day of work.
The queue is not an implementation detail of the network layer. It is the system of record for a window of time, and it has to be designed as storage: durability, ordering, capacity, expiry, and a rule for entries the server will never accept.
Why it matters
A write that has been shown to the user as done cannot be lost, which means it must be on disk before the interface acknowledges it. Holding it in memory is enough to pass every test and lose data the first time the operating system reclaims the process, which on mobile happens routinely and without warning.
The second reason is ordering. Many domain operations are only meaningful in sequence — a job cannot be completed before it is accepted, a refund cannot precede a payment — so the drain is a sequential replay rather than a parallel flush. That makes a single permanently-rejected entry capable of blocking everything behind it, which is the failure that turns a one-hour outage into a week of support tickets.
Implementation patterns
- Durable before acknowledged. Write to the local database inside the same transaction that updates the visible state, then return. If the two are separable, they will separate.
- An operation id generated on the device, persisted before the first send attempt and reused on every retry, so the server collapses duplicates. The lost response is far more common than the lost request on a weak link.
- Classify every entry as superseding or discrete. A location sample is superseded by the next one and can be dropped; a cash collection cannot. The queue has to know which, because that classification is the whole of its overflow policy.
- Cap by bytes and by age, not by entry count, and size the cap from the longest outage the operation plans to survive — a full shift, not a few minutes.
- Quarantine, not blockage. An entry the server rejects with a permanent 4xx moves to a quarantine area, the drain continues, and the entry is surfaced to a human.
- Drain with backpressure awareness: a few entries in flight, honouring
Retry-After, because a thousand devices reconnecting after an outage drain simultaneously.
Industry example
Delivery and field-service platforms on low-end Android handsets, of the kind that characterise last-mile operations in dense South and South-East Asian cities, run this pattern as their core write path: the application commits locally, shows the courier a completed job, and reconciles when the radio returns. The platform primitives are documented rather than bespoke: MQTT 3.1.1 has been an OASIS standard since 2014 and specifies persistent sessions with queued QoS 1 messages for a client that is currently disconnected, and Android's WorkManager (introduced in 2018) supplies the durable retry scheduling a drain loop needs across process death. The common production defect is not the queue's absence but its overflow policy — a queue capped at a few hundred entries, trimmed oldest-first, silently discarding completed jobs during a long network outage, discovered in month-end reconciliation rather than in telemetry.
Failure scenarios
- Memory-only queue: the operating system kills the process to reclaim memory and the day's work is gone, with no error anywhere.
- Oldest-first trim applied to discrete entries: data loss presented as cache eviction.
- A poison entry under strict ordering: one permanently rejected operation stops every later one indefinitely.
- Server-assigned ids: a retry after a lost response creates a second record, which shows up as duplicate charges or double-credited work.
- Unbounded growth: a device offline for three weeks fills its storage and the application cannot start.
- A drain with no pacing: every device in a region reconnects at once and the ingest tier takes the burst that the queue was supposed to smooth.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Store-and-forward on device | Usable in the field; connectivity becomes an optimisation | A second source of truth, a migration path for its schema, and a reconciliation story |
| Direct writes with a spinner | One source of truth and trivial debugging | Unusable wherever the network is not reliable |
The permanent cost is the local schema. Entries written by a version of the app from eight months ago must still be drainable, so the queue's format is a versioned contract with your own past releases, and it cannot be refactored casually.
When not to use it
When the operation cannot be honoured offline. Booking the last part in a shared inventory, authorising a payment, or anything that allocates a scarce resource must not be accepted locally, because the user will be told it succeeded and later told it did not. Queue the intent and show it as pending, or refuse it — never show it as done. For a product whose users are always on reliable connectivity, a retry with a visible error is simpler and better.
Interview question
Q: Your offline queue hits its cap during a six-hour outage. Walk me through what you drop, what you refuse, and what the user sees — and tell me which metric would have warned you a month earlier.
What a strong answer covers: classification of entries into superseding and discrete; dropping only the former; refusing new discrete work with a visible explanation rather than trimming; quarantine for poison entries; and the metric, which is the age of the oldest unsent entry at p99 across the fleet, not queue depth.
Quick check
Quiz: Why must the idempotency key for a queued write be generated on the device and written to disk before the first attempt? Answer: because the response is lost more often than the request, so the retry is certain, and a key assigned by the server on the first attempt is unavailable to the retry.
Flashcard: The on-device queue hits its cap during a long outage. What should it do? — Drop superseding entries oldest-first; stop accepting new discrete work and say so; never trim a committed user intent silently.