Claude Opus 4.6
Anthropic's Opus 4.6 release, positioned around long-running agentic coding, computer-use, and professional workflows.
A comparison of Anthropic and Google frontier model records, focused on available coding and multimodal benchmark evidence.
Anthropic's Opus 4.6 release, positioned around long-running agentic coding, computer-use, and professional workflows.
Google DeepMind's Gemini 3.1 Pro preview, with official benchmark coverage across reasoning, coding, multimodal, tool-use, and long-context tasks.
Claude Opus 4.6 has the higher sourced SWE-bench Verified row in this registry, with its harness caveat shown.
Gemini 3.1 Pro has multiple source-backed rows across GPQA, MMMU-Pro, BrowseComp, Terminal-Bench, and SWE-Bench Pro.
| Model | Benchmark | Score | Source | Source Type | Verified | Notes |
|---|---|---|---|---|---|---|
| Claude Opus 4.6 | SWE-bench Verified | 81.42 | Introducing Claude Opus 4.6 Accessed 2026-06-11 | official | Yes | Anthropic reports this SWE-bench Verified score averaged over 25 trials with a prompt modification. |
| Gemini 3.1 Pro | SWE-bench Verified | 80.6 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | SWE-Bench Verified single-attempt score from Google DeepMind's Gemini 3.1 Pro benchmark table. |
| Gemini 3.1 Pro | SWE-bench Pro | 54.2 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | SWE-Bench Pro public single-attempt score from Google DeepMind's Gemini 3.1 Pro page. |
| Gemini 3.1 Pro | GPQA / GPQA Diamond | 94.3 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | GPQA Diamond no-tools score from Google DeepMind's Gemini 3.1 Pro benchmark table. |
| Gemini 3.1 Pro | MMMU-Pro | 80.5 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | MMMU-Pro no-tools score from Google DeepMind's Gemini 3.1 Pro benchmark table. |
| Gemini 3.1 Pro | BrowseComp | 85.9 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | BrowseComp score with search, Python, and Browse from Google DeepMind's Gemini 3.1 Pro table. |
| Gemini 3.1 Pro | Terminal-Bench 2.x | 68.5 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | Terminal-Bench 2.0 score using the Terminus-2 harness from Google DeepMind's table. |
official
Official Anthropic release page. Use benchmark notes carefully because some reported scores use repeated trials or prompt modifications.
official
Official Google DeepMind model page with benchmark table and methodology link for Gemini 3.1 Pro.