Evaluation data can leak into training

If the questions you test with end up in the training data, the test stops measuring anything.

It happens easily: public benchmarks get scraped, internal evaluation sets get pasted into a fine-tuning job, examples get reused. The model then appears to perform brilliantly on exactly the cases you check, and no better than before on everything else. Keeping evaluation data genuinely separate is unglamorous and is what makes the numbers mean something.

More on AI supply chain