beginner 3 min answer Multiple choice

Users say the assistant's answers lack context, so a team proposes raising chunk size from 400 tokens to 1200 across a 60000-document policy corpus. Retrieval still returns the top 5 chunks. What is the dominant effect of that change?

chunkingretrievalembeddingsprecisioncontext-window
Pick one
Show the full answer Hide the answer

The mechanism

An embedding is a fixed-size vector no matter how much text it represents. A 400-token chunk and a 1200-token chunk both compress to the same 768 or 1024 floats. The longer the chunk, the more the vector drifts towards the average topic of the text it covers, and the weaker its signal for any one specific fact inside it.

A policy paragraph about parental leave eligibility, embedded on its own, sits close to a query about parental leave. That same paragraph embedded as part of a 1200-token section covering leave, expenses and notice periods produces a vector that is close to nothing in particular. Ranking degrades because the thing you are matching against has become blurrier, not because there is less information in the index.

What the change actually buys and costs

The team's instinct is not wrong about the symptom. Retrieved passages that begin mid-sentence, or that reference "the above table" which is no longer present, genuinely produce bad answers. But raising chunk size fixes that by accident and pays for it three times over:

  • Context spend. Top 5 at 400 tokens is 2000 tokens of retrieved material. Top 5 at 1200 is 6000. Most of that extra text is irrelevant to the question, and models attend worse when the relevant sentence is buried among several thousand tokens of near-miss material.
  • Diagnosis gets harder. With larger chunks, a retrieval hit no longer tells you which sentence mattered, so citation and evaluation both get coarser.
  • Fewer independent votes. Five 1200-token chunks are five opinions from the corpus. Five 400-token chunks from five different documents give the answer more corroboration.

Why the other options fail

"Answers improve because every chunk carries surrounding context" is the expected outcome and it does happen for some queries. It is not dominant: the precision loss applies to every query, the context gain only to queries whose answer straddled a boundary. Measure both and the net is usually negative above roughly 600 to 800 tokens for prose.

"Recall improves because there are fewer near-duplicates" inverts the mechanism. Fewer, larger chunks means each retrieval slot is coarser, so the correct passage competes with more unrelated text inside its own chunk. Duplicate suppression is a separate control and does not need bigger chunks.

"Latency roughly triples" is a plausible-sounding infrastructure answer that is simply wrong. Vector dimensionality is unchanged, so each vector is the same size on disk; the index gets smaller because there are a third as many vectors. Query latency barely moves. The cost moves into the generation call instead.

What to do instead

Keep chunks in the 300 to 600 token band and fix the boundary problem directly. Split on document structure rather than character counts so a chunk is a clause, a section or a procedure. Add 10 to 15 percent overlap so a sentence split across a boundary appears whole somewhere. Prepend a short generated context line to each chunk naming the document and section it came from, so the vector carries the location as well as the content. If the answer genuinely needs a whole section, retrieve small and expand: match on the child chunk and pass the parent section to the model.

When bigger chunks win and when this is the wrong call

Legal contracts and regulatory filings where meaning depends on surrounding clauses, and code files where a function is only interpretable with its imports and class context, both justify larger units. The flip condition is measurable: build 60 to 100 labelled questions, measure recall at 5 and answer correctness at both chunk sizes, and let the larger size win only where it actually does. If nobody has measured retrieval separately from answer quality, chunk size is being tuned by anecdote.