gpt.college

SWE-bench Verified

A curated SWE-bench split for resolving real software issues with agentic coding systems.

Last updated: 2026-06-11

What it measures

  • Repository understanding
  • Patch generation
  • Issue resolution

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 Claude Opus 4.6 Anthropic 81.42 Introducing Claude Opus 4.6 Accessed 2026-06-11 official Yes
2 Gemini 3.1 Pro Google 80.6 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes
3 OpenAI GPT-5.2 Thinking OpenAI 80 Introducing GPT-5.2 Accessed 2026-06-11 official Yes
4 Gemini 3 Pro Google 76.2 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes
5 DeepSeek-R1 DeepSeek 49.2 DeepSeek-R1 model card Accessed 2026-06-11 official Yes

Limitations

  • Agent scaffold, tool access, retries, and test budget can materially change results.
  • The paper is static; use the official leaderboard for dynamic results.
  • Record Full, Verified, Lite, Multimodal, and agent scaffold separately.

leaderboard

SWE-bench official leaderboard

The paper is static; use the official leaderboard for dynamic results. Record Full, Verified, Lite, Multimodal, and agent scaffold separately.