What This Research Area Covers
Evaluation research studies how to accurately and fairly measure model capability, spanning benchmark design, methodology for avoiding contamination and gaming, and increasingly, evaluation of complex, multi-step agentic behavior beyond simple question-answering.
Why It Matters
As models improve, existing benchmarks saturate and lose their ability to differentiate between top models, making the ongoing design of genuinely challenging, uncontaminated evaluation methods a critical, if less publicly visible, research area.
Current Research Directions
Designing benchmarks resistant to training data contamination, developing evaluation methods for open-ended and agentic tasks that don't have a single correct answer, and studying the correlation (or lack thereof) between benchmark performance and real-world usefulness are active areas.
Related Pages
Frequently Asked
Why is evaluation harder than it sounds?
Models can be inadvertently trained on benchmark data (contamination), benchmarks saturate once top models score near-perfectly, and many real tasks don't have a single objectively correct answer to score against.
What's benchmark contamination?
When a model has seen benchmark questions during training, inflating its apparent performance beyond genuine capability.
How does this relate to the Benchmarks topic page?
Evaluation covers the research methodology; our Benchmarks page covers the practical, applied side of interpreting benchmark scores.
Where can I see current benchmark standings?
See our Rankings and LLM Leaderboard pages.