Evaluation Set Contamination
also called Benchmark Leakage, Golden Set Overfit
The gradual loss of an evaluation set's ability to predict production quality as its cases leak into training data, into prompts, or into months of iteration against the same examples.
A team runs the same 200-case suite on every prompt change for six months. The score climbs from 0.71 to 0.89, a 25% relative gain, over roughly 180 days. Production complaints do not fall. The suite stopped being a measurement around month two and became a target, and nothing announced the transition.
Contamination arrives by three routes that need different controls.
Into the model. Public benchmarks live in public places — the Hugging Face Hub distributes datasets openly, which is what makes them useful and what makes them crawlable. Anything on the open web can enter a later training corpus, so a model can score well on a benchmark it has effectively seen.
Into the prompt. Cases get copied into the system prompt as examples of good answers, into a few-shot block, or into the retrieval corpus, and the system then answers its own test.
Into the team. The least discussed route: each iteration against a fixed set fits the prompt to that set's idiosyncrasies, and after 100 rounds the score measures familiarity rather than capability.
Why it matters
Evaluation supports one decision: ship or do not ship. A contaminated score does not fail loudly — it reports success, so the decision is taken with confidence and the error surfaces as customer behaviour weeks later.
It also breaks comparison over time: if the set absorbed six months of tuning, this quarter's 0.89 and last quarter's 0.78 are not on the same scale, and the trend line governance relies on is fiction.
Implementation patterns
- Three sets with different rules. A development set to iterate against freely; a validation set used at most weekly; a locked test set touched only for release decisions and never pasted into a prompt or a ticket.
- Refresh from production monthly, labelling real traffic and keeping an anchor subset stable so trends stay comparable.
- Canary strings. A rare unique token in each private case; if a model reproduces it, the set has leaked.
- Watch the gap, not the score. Plot the locked set against a freshly sampled one: a widening gap means the development set has been fitted, and it appears long before customers complain.
- Version the set and record it with every result, or a score compares to nothing.
Industry example
Benchmark contamination is well documented in the public literature on language models: results on widely published test sets drift upward in ways that do not transfer to held-out or privately built variants, which is why serious comparisons now use private or freshly generated sets. The mechanism is mundane — the same openness that makes a public dataset valuable makes it reachable by the next crawl. The internal version needs no crawler: one spreadsheet of cases and a weekly tuning habit produces the same outcome within a quarter.
Failure scenarios
- A model upgrade that scores higher on your suite and worse in production, because it had more exposure to the public cases the suite was built from.
- A regression the suite cannot see, because the prompt was tuned until every case passed and the failure mode lives just outside them.
- A judge prompt containing the test cases, so the grader recognises the answers it grades.
Trade-offs
Locking a test set means fewer cases to iterate on and slower feedback, which teams resist. Refreshing from production costs labelling effort — the real constraint, since labels need the person who knows the right answer — and partly breaks trend continuity. Secrecy costs convenience: contributors cannot see the case their change failed.
When not to use it
A private set built from your own documents and never published is not exposed to the training-data route at all, so the secrecy machinery is unnecessary. The control that still applies is the split and the exposure count, because fitting by iteration happens regardless of who else can see the set.
For a team with 30 cases on a weekly release, formal three-way splits are over-structure. Lock ten cases, iterate on twenty, and refresh from real traffic each month. The rule that scales down: never let the cases you tune against be the cases you decide on.
Interview question
Q: Your eval score has risen from 0.71 to 0.89 over two quarters and the complaint rate is flat. Tell me how you would find out whether the improvement is real.
What a strong answer covers: sampling fresh production cases, labelling them, and scoring both the current and the two-quarter-old prompt on the old set and the fresh one; reading the gap rather than the level; checking for cases leaked into prompts, few-shot blocks or the retrieval corpus; checking the judge's version; counting how many changes each case has gated; then the structural fix of split sets, refresh cadence and set versioning.
Quick check
Quiz: Your suite scores 0.89 and a freshly labelled production sample scores 0.72 on the same rubric. What does the gap mean? — The set no longer represents the task, usually from being iterated against; refresh from production and lock a test split before the next release decision.
Flashcard: Three routes to a contaminated eval set? — Into training data (public sets are crawlable), into the prompt (cases reused as examples or retrieved), into the team (months of tuning on the same cases). Controls: locked split, refresh cadence, canary strings, exposure counts.