beginner 2 min answer Multiple choice

A user asks your internal assistant about a subject the document corpus does not cover at all. The retriever is configured to return the top 5 chunks. What comes back?

embeddingsretrievalabstentionsimilarity-scoreshybrid-retrieval
Pick one
Show the full answer Hide the answer

The mechanism

A nearest-neighbour search answers "which five vectors are closest", and that question always has an answer. The index holds no notion of relevance, only of distance. Ask it about tax law in a corpus of kitchen equipment manuals and it returns the five manuals that happen to be marginally closer to the query vector than the rest, because something must be closest.

Then the generator receives five passages presented as context and a question. Unless it has been told what to do with irrelevant context, it will do what the prompt implies and compose an answer from them — fluently, with citations pointing at real documents that do not support the claim. The confident wrong answer that gets blamed on the model usually starts here.

Why a score threshold does not rescue it by itself

Cosine similarity is a ranking signal, not a probability. The numeric scale is a property of the embedding model and the corpus, so a cutoff of 0.75 means nothing portable: on one model most true matches sit above 0.8, on another they cluster near 0.3, and re-embedding with a new model moves every number at once. A threshold is still useful, but only as a calibrated artifact — fitted on a few hundred labelled query-passage pairs from your own corpus, stored next to the embedding model version, and refitted when either changes.

Why the other options fail

  • An empty result because nothing passes the bar. This is what people assume they configured, and it is what a keyword search would do. A plain approximate-nearest-neighbour query has no bar; you have to add one and tune it.
  • An error from the index for being out of range. Indexes error on dimension mismatch or a missing collection, not on poor matches. Expecting an error here means a whole class of silent failure goes unhandled.
  • Fewer than five chunks because weak matches are discarded. Some retrieval layers do post-filter, which makes this the most dangerous answer: it is true of one configuration and false by default, so the behaviour must be verified rather than assumed.

What to build instead

  1. An abstention path. The prompt must permit "the documents do not cover this", and the eval set must contain unanswerable questions scored on whether the system declines. Without those cases in the set, abstention is never measured and quietly regresses.
  2. A lexical signal beside the vector one. Hybrid retrieval gives you something a vector search cannot: zero keyword hits is positive evidence of absence. Fusing BM25 with dense results is the cheapest reliable no-match detector available.
  3. A grounding check after retrieval. Require the answer to quote the retrieved text, and reject it when the quote is missing.

When not to add an abstention path

For recommendation and "more like this" surfaces, returning the five nearest items is exactly right and an abstention path would be a bug. The distinction is whether a wrong neighbour costs a shrug or a wrong answer in writing.