What This Research Area Covers
This topic covers the specific standardized tests used to measure and compare AI model capability, spanning general knowledge and reasoning, coding, mathematics, and increasingly, long-horizon agentic task completion.
Why It Matters
Benchmarks provide the common yardstick the field uses to track progress and compare models, even though they have well-documented limitations worth understanding before relying on a specific score.
Current Research Directions
Designing benchmarks resistant to saturation and contamination, and developing new benchmarks for emerging capability areas like agentic task completion, are active areas — see our Evaluation topic page for the research methodology behind this.
Related Pages
Frequently Asked
What are the main benchmark categories?
General reasoning, coding, mathematics, and increasingly, long-horizon agentic task completion are the major current categories.
Are benchmark scores reliable?
Useful as one data point, but with well-documented limitations — see our AI Benchmarks Explained page for the practical, applied guidance.
Where can I see current model standings?
See our Rankings and LLM Leaderboard pages.
How is this different from the Evaluation topic page?
Evaluation covers the research methodology behind measuring capability; this page covers the specific benchmarks and standings themselves.