Glossary

Benchmark

A standardized test or dataset used to measure and compare AI model performance — when a paper claims 'state-of-the-art results,' it usually means beating prior models on a specific benchmark.


What it means

A benchmark is a standardized evaluation dataset or task used to measure and compare the performance of AI models. When researchers say a model achieves “state-of-the-art” results, they mean it scores higher than all previously published models on one or more benchmarks.

Examples of widely used benchmarks across different domains:

  • Natural language processing: MMLU (knowledge across 57 subjects), HumanEval (code generation), BioASQ (biomedical question answering)
  • Protein structure prediction: CASP (Critical Assessment of Structure Prediction) — the defining benchmark AlphaFold won in 2020
  • Weather forecasting: ECMWF’s ERA5-based evaluation — used to compare GraphCast, Pangu-Weather, and AIFS
  • Materials science: Matbench — a standardized suite of property prediction tasks for materials
  • Literature extraction: Elicit and similar tools report performance on systematic review datasets

Why it matters for researchers

Benchmark results are not the same as real-world usefulness. A model can top a benchmark and still fail in practice — benchmarks measure specific, narrow tasks under controlled conditions, and the task may not match your actual use case.

Common benchmark limitations to know:

  • Benchmark saturation: When many models score near 100% on a benchmark, it stops being informative — the field typically creates harder benchmarks rather than declaring the problem solved
  • Training data contamination: If a model’s training data included the benchmark’s test set, its score is inflated. This is a known concern for LLM evaluations and is difficult to verify
  • Task mismatch: A protein structure prediction benchmark may evaluate accuracy on well-known folds; your target might be a structurally novel orphan protein where results are much weaker
  • Metric choice: High performance on one metric (e.g., RMSD for structure) does not guarantee good performance on a downstream metric you care about (e.g., binding affinity prediction)

When reading an AI paper: look for the benchmark name, the specific metric, and whether an independent group reproduced the result — not just the number in the abstract.