What it measures
- Abstract reasoning
- Few-shot generalization
- Efficiency-sensitive problem solving
A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.
score; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 77.1 | Gemini 3.1 Pro announcement Accessed 2026-06-12 | official | Yes |
leaderboard
arXiv or technical papers are static; use the ARC Prize leaderboard for scores. Official leaderboard is the current score source. Track eval-set and cost/efficiency notes.
official
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.