A support assistant adds a cross-encoder reranker. Offline NDCG at 10 rises from 0.61 to 0.74 and it ships. Two weeks later, answers drawn from long runbook pages are worse than they were before reranking, while short FAQ answers improved. Nothing errors and no alert fires. What happened?
Show the full answer Hide the answer
The trigger
A cross-encoder scores the query and the passage as one concatenated sequence, and that sequence has a hard maximum length. The MS MARCO cross-encoders distributed with Hugging Face's sentence-transformers library are documented with a maximum sequence length of 512 tokens, and inputs longer than that are truncated rather than rejected.
The runbook chunks are 900 to 1200 tokens. After the query and the special tokens, roughly the first 480 tokens of each passage reach the model. The reranker is scoring the introduction of every runbook and never seeing the procedure, which is where the answer lives. It then ranks a passage that happens to have a relevant-sounding opening above the passage that actually contains the steps.
Why it propagated
Three properties turned a library default into a production regression:
- Truncation is silent. The tokeniser truncates and returns a score. There is no error, no warning and no field in the response saying the input was cut, so nothing downstream can notice.
- The eval set did not contain the failing shape. NDCG at 10 rose because the labelled questions were built from the FAQ, whose chunks are 150 to 300 tokens and fit comfortably. An offline metric measured on the wrong distribution is worse than no metric, because it authorises the change.
- Reranking is a reordering, so the damage is invisible in retrieval metrics. The correct passage is still in the candidate set. It has simply been pushed from position 2 to position 14, below the cut passed to the model, so the generation step never sees it.
Why detection lagged
Every operational signal stayed flat: latency, error rate, token usage, retrieval recall at 100. The only moved signal was answer quality on a minority of traffic, which nobody was measuring per document type. The user-visible symptom was a confident answer citing the right document and giving the wrong procedure — which reads as a model problem, so the first two weeks of investigation went to prompts.
The structural fix, and the tempting local one
The tempting fix is to raise the reranker's max length. It does not work: these models have a positional limit baked into their pretrained encoder, so raising it either errors or produces untrained behaviour.
The structural fix is to make chunk length a contract with the reranker. Cap indexed chunks at roughly 400 tokens of passage text so query plus passage fits inside 512 with headroom, or move to a reranker trained for longer inputs, or score long passages in overlapping windows and take the maximum. Then make truncation observable: log the tokenised input length and a truncated boolean on every rerank call, and alert when the truncation rate over any document source exceeds 1%.
The general lesson
Every model in a retrieval pipeline has an input limit, and the limits are not the same. The embedding model, the reranker and the generator each truncate differently, and each does it quietly. The pipeline is only correct when chunk sizing is chosen against the tightest of them, not against the generator's context window.
When not to rerank at all, and when it is the wrong investment
If recall at 10 is already close to recall at 100, the candidate set is not the problem and reranking buys nothing while adding a cross-encoder to the request path. Measure that ratio before adding the stage. When the gap is under a few points, choose query rewriting or better generation instead: a cross-encoder costs 30 to 80 ms of p95 on top of every query and adds a model to operate, and it can only reorder what the first stage already found.