What it measures
- Advanced math
- Research-level problem solving
- Proof and derivation robustness
An advanced mathematics benchmark curated around difficult research-style math problems.
% solved; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | OpenAI GPT-5.4 | OpenAI | 47.6 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes |
| 2 | OpenAI GPT-5.2 Thinking | OpenAI | 40.3 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes |
leaderboard
The paper is static; Epoch AI's page is the current hub. Epoch page aggregates tiers 1-4 and open-problem notes.
paper
Use for benchmark definition, not necessarily for latest model scores.
official
Use for dataset, code, or implementation details; score freshness depends on the benchmark.