GPQA / GPQA Diamond
A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.
Signals for synthesis, factual work, long answers, RAG, and source-heavy analysis.
| Rank | Model | Provider | Score | Coverage | Source | Source Type | Verified | Last Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek-R1 | DeepSeek | 79.3 | 3/9 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes | 2026-06-11 |
| Model | Provider | Score | Coverage | Freshest Source | Last Updated |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 90.1 | 2/9 | Gemini 3.1 Pro model page Accessed 2026-06-11 | 2026-06-11 | |
| OpenAI GPT-5.4 | OpenAI | 87.8 | 2/9 | Introducing GPT-5.4 Accessed 2026-06-11 | 2026-06-11 |
| OpenAI GPT-5.2 Thinking | OpenAI | 79.1 | 2/9 | Introducing GPT-5.2 Accessed 2026-06-11 | 2026-06-11 |
| Gemini 3 Pro | 75.6 | 2/9 | Gemini 3.1 Pro model page Accessed 2026-06-11 | 2026-06-11 |
A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.
A harder MMLU-style benchmark for broad knowledge and reasoning.
A machine-learning engineering benchmark based on Kaggle-style competition tasks.
A browsing benchmark for hard-to-find information and deep web research tasks.
A long-context understanding benchmark with a project leaderboard.
A long-context benchmark suite focused on practical application-style tasks.
A chart and figure understanding benchmark for scientific visual reasoning.
A factuality benchmark for short-answer factual knowledge questions.
A RAG and factual multi-hop question answering benchmark.