SWE-bench Verified
A curated SWE-bench split for resolving real software issues with agentic coding systems.
Signals for tool use, browsing, GUI control, terminal operation, and multi-step autonomous workflows.
| Rank | Model | Provider | Score | Coverage | Source | Source Type | Verified | Last Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 72.3 | 4/10 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | 2026-06-11 | |
| 2 | OpenAI GPT-5.2 Thinking | OpenAI | 67.1 | 3/10 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | 2026-06-11 |
| Model | Provider | Score | Coverage | Freshest Source | Last Updated |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 81.4 | 1/10 | Introducing Claude Opus 4.6 Accessed 2026-06-11 | 2026-06-11 |
| OpenAI GPT-5.4 | OpenAI | 70.2 | 2/10 | Introducing GPT-5.4 Accessed 2026-06-11 | 2026-06-11 |
| Gemini 3 Pro | 67.7 | 2/10 | Gemini 3.1 Pro model page Accessed 2026-06-11 | 2026-06-11 | |
| DeepSeek-R1 | DeepSeek | 49.2 | 1/10 | DeepSeek-R1 model card Accessed 2026-06-11 | 2026-06-11 |
A curated SWE-bench split for resolving real software issues with agentic coding systems.
A professional software-engineering agent benchmark with public and private leaderboard splits.
A terminal-environment agent benchmark for command-line task completion.
A machine-learning engineering benchmark based on Kaggle-style competition tasks.
A browsing benchmark for hard-to-find information and deep web research tasks.
A tool-use benchmark family for business-process and customer-service style agent tasks.
A computer-use benchmark for operating-system and GUI agents.
A web-agent benchmark for completing tasks in realistic self-hosted web environments.
A visual web-agent benchmark extending WebArena with image-grounded web tasks.
A terminal-agent benchmark candidate that needs source confirmation.