gpt.college

Overall Ranking

A cautious editorial rollup of the selected benchmark set. It is not an absolute model quality score.

Main rankings require at least 3 of 25 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated
1 Gemini 3.1 Pro Google 73.6 8/25 Gemini 3.1 Pro announcement Accessed 2026-06-12 official Yes 2026-06-12
2 Gemini 3 Pro Google 69.2 5/25 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12
3 OpenAI GPT-5.2 Thinking OpenAI 68.9 6/25 Introducing GPT-5.2 Accessed 2026-06-11 official Yes 2026-06-11
4 OpenAI GPT-5.4 OpenAI 66.4 6/25 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12
5 DeepSeek-R1 DeepSeek 59.1 5/25 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
Claude Opus 4.6 Anthropic 58.8 2/25 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12

score

ARC-AGI-2

A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.

% accuracy

Humanity's Last Exam (HLE)

A broad expert-level benchmark intended to test difficult questions across many domains and formats.

% solved

FrontierMath

An advanced mathematics benchmark curated around difficult research-style math problems.

% accuracy

GPQA / GPQA Diamond

A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.

% accuracy

MMLU-Pro

A harder MMLU-style benchmark for broad knowledge and reasoning.

% resolved

SWE-bench Verified

A curated SWE-bench split for resolving real software issues with agentic coding systems.

% resolved

SWE-bench Pro

A professional software-engineering agent benchmark with public and private leaderboard splits.

% success

Terminal-Bench 2.x

A terminal-environment agent benchmark for command-line task completion.

medal or score

MLE-bench

A machine-learning engineering benchmark based on Kaggle-style competition tasks.

% accuracy

BrowseComp

A browsing benchmark for hard-to-find information and deep web research tasks.

% success

tau-bench / tau2-bench

A tool-use benchmark family for business-process and customer-service style agent tasks.

% success

WebArena

A web-agent benchmark for completing tasks in realistic self-hosted web environments.

% success

VisualWebArena

A visual web-agent benchmark extending WebArena with image-grounded web tasks.

% accuracy

LongBench v2

A long-context understanding benchmark with a project leaderboard.

% accuracy

RULER

A synthetic long-context stress benchmark from NVIDIA.

score

HELMET

A long-context benchmark suite focused on practical application-style tasks.

% accuracy

MMMU-Pro

A harder multimodal academic benchmark derived from MMMU-style tasks.

% accuracy

Video-MME

A video understanding benchmark for multimodal models.

% accuracy

MathVista

A visual math reasoning benchmark for multimodal models.

% accuracy

CharXiv

A chart and figure understanding benchmark for scientific visual reasoning.

% correct

SimpleQA Verified

A factuality benchmark for short-answer factual knowledge questions.

% accuracy

FRAMES

A RAG and factual multi-hop question answering benchmark.

score

GraphWalks

A graph reasoning benchmark candidate that still needs a stable authoritative source before launch.