gpt.college

Research Ranking

Signals for synthesis, factual work, long answers, RAG, and source-heavy analysis.

Main rankings require at least 3 of 9 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated
1 DeepSeek-R1 DeepSeek 79.3 3/9 DeepSeek-R1 model card Accessed 2026-06-11 official Yes 2026-06-11

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
Gemini 3.1 Pro Google 90.1 2/9 Gemini 3.1 Pro model page Accessed 2026-06-11 2026-06-11
OpenAI GPT-5.4 OpenAI 87.8 2/9 Introducing GPT-5.4 Accessed 2026-06-11 2026-06-11
OpenAI GPT-5.2 Thinking OpenAI 79.1 2/9 Introducing GPT-5.2 Accessed 2026-06-11 2026-06-11
Gemini 3 Pro Google 75.6 2/9 Gemini 3.1 Pro model page Accessed 2026-06-11 2026-06-11

% accuracy

GPQA / GPQA Diamond

A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.

% accuracy

MMLU-Pro

A harder MMLU-style benchmark for broad knowledge and reasoning.

medal or score

MLE-bench

A machine-learning engineering benchmark based on Kaggle-style competition tasks.

% accuracy

BrowseComp

A browsing benchmark for hard-to-find information and deep web research tasks.

% accuracy

LongBench v2

A long-context understanding benchmark with a project leaderboard.

score

HELMET

A long-context benchmark suite focused on practical application-style tasks.

% accuracy

CharXiv

A chart and figure understanding benchmark for scientific visual reasoning.

% correct

SimpleQA Verified

A factuality benchmark for short-answer factual knowledge questions.

% accuracy

FRAMES

A RAG and factual multi-hop question answering benchmark.