Strengths
- Strong GPQA Diamond and MMMU-Pro scores
- Broad multimodal model coverage
A Gemini 3 Pro baseline included in Google DeepMind's Gemini 3.1 Pro comparison table.
| Benchmark | Score | Source | Source Type | Verified | Notes |
|---|---|---|---|---|---|
| SWE-bench Verified | 76.2 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | SWE-Bench Verified comparison-row score for Gemini 3 Pro on the Gemini 3.1 Pro page. |
| GPQA / GPQA Diamond | 91.9 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | GPQA Diamond comparison-row score for Gemini 3 Pro on the Gemini 3.1 Pro page. |
| MMMU-Pro | 81 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | MMMU-Pro no-tools comparison-row score for Gemini 3 Pro on the Gemini 3.1 Pro page. |
| BrowseComp | 59.2 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | BrowseComp comparison-row score for Gemini 3 Pro on the Gemini 3.1 Pro page. |
| Humanity's Last Exam (HLE) | 37.72 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | Scale Labs HLE Text Only score for gemini-3-pro-preview, covering the 86% text-only HLE subset. |
official
Use for model availability and context notes; verify before publishing claims.
official
Official Google DeepMind model page with benchmark table and methodology link for Gemini 3.1 Pro.