gpt.college

Reasoning Ranking

Signals for abstract reasoning, hard question answering, math, and multi-step problem solving.

Main rankings require at least 3 of 6 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated
1 Gemini 3.1 Pro Google 72.9 3/6 Gemini 3.1 Pro announcement Accessed 2026-06-12 official Yes 2026-06-12
2 OpenAI GPT-5.4 OpenAI 59 3/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12
3 DeepSeek-R1 DeepSeek 54.7 3/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
OpenAI GPT-5.2 Thinking OpenAI 66.3 2/6 Introducing GPT-5.2 Accessed 2026-06-11 2026-06-11
Gemini 3 Pro Google 64.8 2/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12
Claude Opus 4.6 Anthropic 36.2 1/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12

score

ARC-AGI-2

A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.

% accuracy

Humanity's Last Exam (HLE)

A broad expert-level benchmark intended to test difficult questions across many domains and formats.

% solved

FrontierMath

An advanced mathematics benchmark curated around difficult research-style math problems.

% accuracy

GPQA / GPQA Diamond

A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.

% accuracy

MMLU-Pro

A harder MMLU-style benchmark for broad knowledge and reasoning.

score

GraphWalks

A graph reasoning benchmark candidate that still needs a stable authoritative source before launch.