gpt.college

Knowledge Ranking

Signals for expert knowledge, broad academic coverage, and factual accuracy.

Main rankings require at least 3 of 6 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated
1 Gemini 3.1 Pro Google 74 3/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12
2 Gemini 3 Pro Google 70.2 3/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12
3 OpenAI GPT-5.4 OpenAI 70.2 3/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12
4 DeepSeek-R1 DeepSeek 61.6 4/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes 2026-06-12

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
OpenAI GPT-5.2 Thinking OpenAI 86 2/6 Introducing GPT-5.2 Accessed 2026-06-11 2026-06-11
Claude Opus 4.6 Anthropic 36.2 1/6 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12

% accuracy

Humanity's Last Exam (HLE)

A broad expert-level benchmark intended to test difficult questions across many domains and formats.

% accuracy

GPQA / GPQA Diamond

A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.

% accuracy

MMLU-Pro

A harder MMLU-style benchmark for broad knowledge and reasoning.

% accuracy

MMMU-Pro

A harder multimodal academic benchmark derived from MMMU-style tasks.

% correct

SimpleQA Verified

A factuality benchmark for short-answer factual knowledge questions.

% accuracy

FRAMES

A RAG and factual multi-hop question answering benchmark.