Practice · Fast-moving · Intermediate
AI Benchmark
A standard task and dataset used to compare models on a shared yardstick.
What AI Benchmark is
Benchmarks range from narrow academic sets to broad suites covering reasoning, coding, maths and multilingual ability. They coordinate the field but also distort it once optimisation targets them.
How it works
Models are run under a documented protocol and scored automatically or by human or model graders. Credible reporting states the prompting method, the number of attempts and whether the test set may have leaked into training.
Why it matters
Benchmarks drive purchasing and research decisions, so knowing their limits — contamination, saturation, narrow coverage — matters as much as knowing the scores.
Common uses
- →Model selection for a product
- →Tracking capability progress
- →Research publication comparison
Strengths
- ✓Comparable across labs
- ✓Fast signal
Watch for
- ✓Contamination inflates results
- ✓High scores may not transfer to your task
Continue exploring
More in this collection
Browse all AI Concepts