A web-indexing pipeline of the kind Baidu operates has six queue-connected stages. Roughly one document in two million makes the parse stage abort; the queue redelivers it, so that partition crash-loops and 40 minutes of documents sit behind it while the other partitions drain normally. Review the pipeline. Which single change would you make first?
Show the full answer Hide the answer
What is actually required
A staged pipeline needs one invariant: every record ends in exactly one terminal state — processed, or set aside somewhere with a reason and a count. This pipeline has no third state, so a record the parser cannot handle has only two options, and it takes the partition with it.
Size the problem before choosing: at one failure in two million and a few hundred million documents a day, that is on the order of a hundred poison records a day. Three parse attempts each costing a second or two of CPU is nothing. What is expensive is 40 minutes of head-of-line blocking per poison record, which at the pipeline's throughput is the real loss, and it is loss nobody is being paged for.
The one change that matters
A bounded attempt count plus a durable quarantine with an observable depth. Three mechanical requirements:
- Attempt count on the record, not in the consumer. A redelivery counter held in memory resets on restart, which is exactly when a crash-looping record comes back around.
- Quarantine carries the payload, the stage, the exception and the attempt history, so the record can be replayed after a parser fix rather than re-crawled.
- Alert on a step change, not on a non-zero depth. A threshold of zero pages every night. A rule such as "quarantine arrivals in the last hour above five times the trailing seven-day median" is what catches the parser regression that starts rejecting an entire site.
Why the other options fail
- Log the error and drop. It unblocks the partition and converts a loud failure into an invisible one. In a search corpus a dropped document is a site that quietly stops appearing, and nobody discovers it from a log line. The honest version of dropping is dropping into somewhere with a count, which is quarantine.
- Validate at ingest. Appealing, and incomplete: to know that this document breaks the parser you have to run the parser. A validator strong enough is the parser, so the crash moves to stage zero and the queue in front of it. Cheap structural checks at ingest are still worth having — they move the common rejects earlier — but they do not remove the need for a terminal state.
- Longer backoff. The record still blocks its partition, just at a politer interval. It makes the stall last longer and look less like an incident.
What I would leave alone
The six stages and the queues between them. Independent retry and independent scaling per stage are what this shape buys, and nothing about this failure argues for collapsing them. Resist the instinct to answer a poison-record problem with a topology change.
When this is the wrong answer
When continuing is worse than stopping. For a keyed stream of financial events where per-key ordering is part of the contract, quietly setting a record aside and processing the next one for the same key produces a state that is wrong rather than late. There the right behaviour is to halt that partition deliberately and page a human — a poison-pill stop, with the stall as the intended signal. The deciding question is whether a gap in the output is tolerable: for an index, yes; for a ledger, no.