ARC-AGI-2
A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.
A cautious editorial rollup of the selected benchmark set. It is not an absolute model quality score.
| Rank | Model | Provider | Score | Coverage | Source | Source Type | Verified | Last Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 73.6 | 8/25 | Gemini 3.1 Pro announcement Accessed 2026-06-12 | official | Yes | 2026-06-12 | |
| 2 | Gemini 3 Pro | 69.2 | 5/25 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | 2026-06-12 | |
| 3 | OpenAI GPT-5.2 Thinking | OpenAI | 68.9 | 6/25 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | 2026-06-11 |
| 4 | OpenAI GPT-5.4 | OpenAI | 66.4 | 6/25 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | 2026-06-12 |
| 5 | DeepSeek-R1 | DeepSeek | 59.1 | 5/25 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | 2026-06-12 |
| Model | Provider | Score | Coverage | Freshest Source | Last Updated |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 58.8 | 2/25 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 |
A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.
A broad expert-level benchmark intended to test difficult questions across many domains and formats.
An advanced mathematics benchmark curated around difficult research-style math problems.
A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.
A harder MMLU-style benchmark for broad knowledge and reasoning.
A curated SWE-bench split for resolving real software issues with agentic coding systems.
A professional software-engineering agent benchmark with public and private leaderboard splits.
A terminal-environment agent benchmark for command-line task completion.
A machine-learning engineering benchmark based on Kaggle-style competition tasks.
A browsing benchmark for hard-to-find information and deep web research tasks.
A tool-use benchmark family for business-process and customer-service style agent tasks.
A computer-use benchmark for operating-system and GUI agents.
A web-agent benchmark for completing tasks in realistic self-hosted web environments.
A visual web-agent benchmark extending WebArena with image-grounded web tasks.
A long-context understanding benchmark with a project leaderboard.
A synthetic long-context stress benchmark from NVIDIA.
A long-context benchmark suite focused on practical application-style tasks.
A harder multimodal academic benchmark derived from MMMU-style tasks.
A video understanding benchmark for multimodal models.
A visual math reasoning benchmark for multimodal models.
A chart and figure understanding benchmark for scientific visual reasoning.
A factuality benchmark for short-answer factual knowledge questions.
A RAG and factual multi-hop question answering benchmark.
A graph reasoning benchmark candidate that still needs a stable authoritative source before launch.
A terminal-agent benchmark candidate that needs source confirmation.