What it measures
- Scientific reasoning
- Expert question answering
- Hard multiple-choice evaluation
A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 94.3 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | |
| 2 | OpenAI GPT-5.4 | OpenAI | 92.8 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes |
| 3 | OpenAI GPT-5.2 Thinking | OpenAI | 92.4 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes |
| 4 | Gemini 3 Pro | 91.9 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | |
| 5 | DeepSeek-R1 | DeepSeek | 71.5 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes |
unknown
The arXiv paper defines the benchmark and original baselines; latest scores usually come from model releases. No single stable dynamic official leaderboard was selected; cite each model-release result separately.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.