Skip to content
AI360Xpert
Glossary
Definition

Benchmark Contamination

A catastrophic evaluation failure where the exact test questions used to evaluate a language model were accidentally included in its massive training corpus.

Think of It Like This

Like a student stealing the final exam answer key a week before the test; their perfect score proves they can memorize, not that they actually learned the material.

When a model achieves superhuman scores on a reasoning test, it is often because it simply memorized the answers during pre-training rather than actually learning to reason. This forces researchers to constantly invent new, secret benchmarks or use dynamic evaluation methods to get a true measure of a model's zero-shot generalization capabilities.