What it measures
- Knowledge breadth
- Reasoning under harder answer choices
- Academic subject coverage
A harder MMLU-style benchmark for broad knowledge and reasoning.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | DeepSeek-R1 | DeepSeek | 84 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes |
leaderboard
The arXiv paper is static; the leaderboard and GitHub may update. HF Space is the score entry point; GitHub is the implementation entry point.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.