Evaluation data can leak into training
If the questions you test with end up in the training data, the test stops measuring anything.
It happens easily: public benchmarks get scraped, internal evaluation sets get pasted into a fine-tuning job, examples get reused. The model then appears to perform brilliantly on exactly the cases you check, and no better than before on everything else. Keeping evaluation data genuinely separate is unglamorous and is what makes the numbers mean something.
More on AI supply chain
- A model file is executable trust in another formIt looks like data until you open it
- Dataset provenance matters for security and governanceWhere did this batch come from?
- Model version changes can be security changesOne plate swapped inside
- Third-party AI APIs extend the data boundaryThe fence moves with the call
- Fine-tuning credentials are production credentialsThe bench feeds the floor
- Open models shift responsibility toward the operatorThe engine comes with the engine room
