gpt.college

ARC-AGI-2

A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.

Last updated: 2026-06-11

What it measures

  • Abstract reasoning
  • Few-shot generalization
  • Efficiency-sensitive problem solving

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 Gemini 3.1 Pro Google 77.1 Gemini 3.1 Pro announcement Accessed 2026-06-12 official Yes

Limitations

  • Scores should be read with the stated ARC Prize evaluation and efficiency rules.
  • arXiv or technical papers are static; use the ARC Prize leaderboard for scores.
  • Official leaderboard is the current score source. Track eval-set and cost/efficiency notes.

leaderboard

ARC Prize official leaderboard

arXiv or technical papers are static; use the ARC Prize leaderboard for scores. Official leaderboard is the current score source. Track eval-set and cost/efficiency notes.

third-party

ARC-AGI-2 dataset or code

Use for dataset, code, or implementation details; score freshness depends on the benchmark.