gpt.college

GPQA / GPQA Diamond

A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.

Last updated: 2026-06-11

What it measures

  • Scientific reasoning
  • Expert question answering
  • Hard multiple-choice evaluation

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 Gemini 3.1 Pro Google 94.3 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes
2 OpenAI GPT-5.4 OpenAI 92.8 Introducing GPT-5.4 Accessed 2026-06-11 official Yes
3 OpenAI GPT-5.2 Thinking OpenAI 92.4 Introducing GPT-5.2 Accessed 2026-06-11 official Yes
4 Gemini 3 Pro Google 91.9 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes
5 DeepSeek-R1 DeepSeek 71.5 DeepSeek-R1 model card Accessed 2026-06-11 official Yes

Limitations

  • Separate GPQA and GPQA Diamond variants in score records.
  • The arXiv paper defines the benchmark and original baselines; latest scores usually come from model releases.
  • No single stable dynamic official leaderboard was selected; cite each model-release result separately.

unknown

Model release pages plus GPQA paper and repository

The arXiv paper defines the benchmark and original baselines; latest scores usually come from model releases. No single stable dynamic official leaderboard was selected; cite each model-release result separately.