What it measures
- Complex engineering tasks
- Agentic coding
- Repository-scale issue resolution
A professional software-engineering agent benchmark with public and private leaderboard splits.
% resolved; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | OpenAI GPT-5.4 | OpenAI | 57.7 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes |
| 2 | OpenAI GPT-5.2 Thinking | OpenAI | 55.6 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes |
| 3 | Gemini 3.1 Pro | 54.2 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes |
leaderboard
The paper is static; Scale's leaderboard is dynamic. Separate public and private splits in result records.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.