gpt.college

Claude Opus 4.6 vs Gemini 3.1 Pro

A comparison of Anthropic and Google frontier model records, focused on available coding and multimodal benchmark evidence.

Last updated: 2026-06-11

Claude Opus 4.6

Anthropic's Opus 4.6 release, positioned around long-running agentic coding, computer-use, and professional workflows.

codingreasoningwritingresearchagentic

Gemini 3.1 Pro

Google DeepMind's Gemini 3.1 Pro preview, with official benchmark coverage across reasoning, coding, multimodal, tool-use, and long-context tasks.

codingreasoningmultimodalresearchlong-contextagentic

Use Case Fit

Coding evidence

Claude Opus 4.6

Claude Opus 4.6 has the higher sourced SWE-bench Verified row in this registry, with its harness caveat shown.

Broad multimodal and agentic coverage

Gemini 3.1 Pro

Gemini 3.1 Pro has multiple source-backed rows across GPQA, MMMU-Pro, BrowseComp, Terminal-Bench, and SWE-Bench Pro.

Benchmark Signals

Model Benchmark Score Source Source Type Verified Notes
Claude Opus 4.6 SWE-bench Verified 81.42 Introducing Claude Opus 4.6 Accessed 2026-06-11 official Yes Anthropic reports this SWE-bench Verified score averaged over 25 trials with a prompt modification.
Gemini 3.1 Pro SWE-bench Verified 80.6 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes SWE-Bench Verified single-attempt score from Google DeepMind's Gemini 3.1 Pro benchmark table.
Gemini 3.1 Pro SWE-bench Pro 54.2 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes SWE-Bench Pro public single-attempt score from Google DeepMind's Gemini 3.1 Pro page.
Gemini 3.1 Pro GPQA / GPQA Diamond 94.3 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes GPQA Diamond no-tools score from Google DeepMind's Gemini 3.1 Pro benchmark table.
Gemini 3.1 Pro MMMU-Pro 80.5 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes MMMU-Pro no-tools score from Google DeepMind's Gemini 3.1 Pro benchmark table.
Gemini 3.1 Pro BrowseComp 85.9 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes BrowseComp score with search, Python, and Browse from Google DeepMind's Gemini 3.1 Pro table.
Gemini 3.1 Pro Terminal-Bench 2.x 68.5 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes Terminal-Bench 2.0 score using the Terminus-2 harness from Google DeepMind's table.

official

Introducing Claude Opus 4.6

Official Anthropic release page. Use benchmark notes carefully because some reported scores use repeated trials or prompt modifications.

official

Gemini 3.1 Pro model page

Official Google DeepMind model page with benchmark table and methodology link for Gemini 3.1 Pro.