Benchmark contamination occurs when evaluation questions, answers, or close variants appear in a model's training or tuning data. The model can then score well through prior exposure rather than general capability. Contamination weakens comparisons, especially for public benchmarks whose contents have circulated widely across the web.
How does contamination distort an evaluation?
The held-out test no longer measures unseen performance. A higher score may reflect memorization rather than transfer to new tasks. Research on multilingual benchmarks describes contamination as test data appearing in pretraining or post-training corpora. Exact matching catches only the simplest leakage.
How can teams reduce the risk?
Keep the eval dataset private where possible, create fresh task-specific cases, check cross-split overlap, and separate evaluation from training pipelines. Confirm model choices on real workflow outcomes, not one public leaderboard. Synthetic data can create variants but may reproduce exposed material. An AI evaluation harness should record dataset versions, model versions, and any known exposure.
Frequently asked questions
Is benchmark contamination the same as overfitting?
It is a specific leakage problem where test material enters training, while overfitting more broadly means learning patterns that do not generalize.
Can exact duplicate checks prove a benchmark is clean?
No. Paraphrases, translations, derived examples, and inaccessible training corpora make contamination difficult to rule out.
No advertising or tracking cookies, and our visitor counts are anonymous. The Cal.com booking widget loads only if you allow it. Privacy Policy.
The page itself, anything our host sets to serve and secure it, and the anonymous visitor count. Always on, and none of it stores anything on your device.
The Cal.com booking widget. Left off, a booking link opens the booking page instead of a popup, so you can still book a call.