Benchmarks measure selected tasks, not universal capability
A benchmark score tells you how a model did on a particular set of questions, chosen by somebody, at a point in time.
It does not transfer cleanly to your use case, and the questions leak into training data over time, which inflates later scores without anything improving. Benchmarks are useful for rough comparison and routinely used as if they were a guarantee. The only evaluation that tells you much about your application is one built from your own tasks.
More on AI and LLMs
- An LLM predicts plausible continuations, not verified truthIt only checks the shape
- Context is temporary working material, not permanent knowledgeThe board gets wiped
- Retrieval adds documents, not guaranteed correctnessThe filter slot is empty
- Temperature changes variation, not factualityThe dial only sets the spread
- System prompts are instructions, not a security boundaryA sign, with no fence
- LLM output is untrusted input downstreamIt comes in round the back
