What it measures
- Repository understanding
- Patch generation
- Issue resolution
A curated SWE-bench split for resolving real software issues with agentic coding systems.
% resolved; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 81.42 | Introducing Claude Opus 4.6 Accessed 2026-06-11 | official | Yes |
| 2 | Gemini 3.1 Pro | 80.6 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | |
| 3 | OpenAI GPT-5.2 Thinking | OpenAI | 80 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes |
| 4 | Gemini 3 Pro | 76.2 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes | |
| 5 | DeepSeek-R1 | DeepSeek | 49.2 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes |
leaderboard
The paper is static; use the official leaderboard for dynamic results. Record Full, Verified, Lite, Multimodal, and agent scaffold separately.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.