Benchmarks measure selected tasks, not universal capability

A benchmark score tells you how a model did on a particular set of questions, chosen by somebody, at a point in time.

It does not transfer cleanly to your use case, and the questions leak into training data over time, which inflates later scores without anything improving. Benchmarks are useful for rough comparison and routinely used as if they were a guarantee. The only evaluation that tells you much about your application is one built from your own tasks.

More on AI and LLMs