ARC-AGI-2
A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.
Signals for abstract reasoning, hard question answering, math, and multi-step problem solving.
| Rank | Model | Provider | Score | Coverage | Source | Source Type | Verified | Last Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 72.9 | 3/6 | Gemini 3.1 Pro announcement Accessed 2026-06-12 | official | Yes | 2026-06-12 | |
| 2 | OpenAI GPT-5.4 | OpenAI | 59 | 3/6 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | 2026-06-12 |
| 3 | DeepSeek-R1 | DeepSeek | 54.7 | 3/6 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | 2026-06-12 |
| Model | Provider | Score | Coverage | Freshest Source | Last Updated |
|---|---|---|---|---|---|
| OpenAI GPT-5.2 Thinking | OpenAI | 66.3 | 2/6 | Introducing GPT-5.2 Accessed 2026-06-11 | 2026-06-11 |
| Gemini 3 Pro | 64.8 | 2/6 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 | |
| Claude Opus 4.6 | Anthropic | 36.2 | 1/6 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 |
A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.
A broad expert-level benchmark intended to test difficult questions across many domains and formats.
An advanced mathematics benchmark curated around difficult research-style math problems.
A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.
A harder MMLU-style benchmark for broad knowledge and reasoning.
A graph reasoning benchmark candidate that still needs a stable authoritative source before launch.