Also called: eval, leaderboard
A fixed set of tasks used to score models against each other. MMLU, SWE-bench, GPQA and friends.
Why it mattersUseful for rough ranking. A benchmark win rarely predicts whether a model is good at your job.
See also Evals