Visibility Timeout
also called Message Lease, Invisibility Window
The period a claimed message is hidden from other consumers - too short and work is duplicated, too long and a crashed worker stalls the item.
When a worker claims a message from a queue, the message is not deleted; it is hidden for a defined period. If the worker acknowledges before the period expires, the message is removed. If it does not — because the worker crashed, stalled, or is simply slow — the message becomes visible again and another worker claims it.
This is what makes competing consumers recoverable without a coordinator: no worker holds an exclusive lock, and a lost worker costs one visibility period rather than a stuck item forever.
The tension
Too short and a slow-but-healthy worker loses its claim mid-processing. Another worker starts the same job, and both complete — the most common cause of duplicate processing, and it looks like a queue defect when it is a configuration one.
Too long and a crashed worker's item waits the full period before anyone retries, which for a latency-sensitive workload is a stall.
There is no setting that removes the tension, which is why the durable answer is idempotency rather than tuning — exactly as with distributed locks and leases.
Implementation patterns
- Set it from measured processing time, at a high percentile with margin — not from the average.
- Extend it during long work. A worker processing a large job periodically extends its claim, so the timeout can be short for the common case without penalising the tail.
- Idempotent processing, always. Delivery is at-least-once, so every consumer will occasionally process the same message twice. This is the actual correctness mechanism.
- Acknowledge after completion, never on receipt, or a crash loses the work silently.
- A retry ceiling with a dead-letter path, so one item that consistently fails does not consume workers forever.
- Monitor redelivery counts. A rising redelivery rate is the signature of either a timeout set too short or a poison message, and distinguishing them is the first diagnostic step.
Industry example
Media rendering pipelines make the parameters concrete: jobs range from seconds to many minutes, so a single timeout value is wrong for most of the distribution. The workable design combines a short default with extension during processing, plus deterministic output locations and conditional completion records so a duplicate render costs compute rather than correctness.
The same structure appears wherever workers process variable-length work from a shared queue — video transcoding, document conversion, batch imports, report generation — and the recurring failure is a timeout tuned to the median with a long tail of duplicated work nobody attributes correctly.
Failure scenarios
- Timeout shorter than the tail of processing time, producing steady duplicate work.
- Acknowledging on receipt, so crashes lose work silently.
- No retry ceiling, so a poison message permanently occupies a worker slot and effective capacity quietly falls.
- Extension implemented on the same thread as the work, so a long operation blocks its own extension and loses the claim it is actively using — the same defect as lease renewal on a busy thread.
- Relying on the timeout for correctness rather than on idempotency.
Trade-offs
Short timeouts recover quickly from crashes and duplicate more; long timeouts duplicate less and stall longer. Extension adds protocol complexity and removes most of the tension for variable-length work.
The framing that resolves it: the visibility timeout is an optimisation that limits wasted work, and idempotency is what makes the system correct. A design whose correctness depends on the timeout being right is a design that will be wrong under load.
Interview question
"Your workers occasionally process the same job twice and you have not changed anything. Give me three explanations, tell me which is most likely, and tell me what makes the system correct regardless."