Strengths
- Strong GPQA and FrontierMath performance
- Useful cross-category benchmark coverage
- Explicit API model naming
A reasoning-focused GPT-5.2 configuration with detailed official benchmark tables across coding, academic, vision, and tool-use tasks.
| Benchmark | Score | Source | Source Type | Verified | Notes |
|---|---|---|---|---|---|
| SWE-bench Verified | 80 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | SWE-bench Verified score reported by OpenAI for GPT-5.2 Thinking. |
| SWE-bench Pro | 55.6 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | SWE-Bench Pro public score reported by OpenAI for GPT-5.2 Thinking. |
| GPQA / GPQA Diamond | 92.4 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | GPQA Diamond no-tools score from OpenAI's GPT-5.2 academic benchmark table. |
| FrontierMath | 40.3 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | FrontierMath Tier 1-3 with Python score from OpenAI's GPT-5.2 release page. |
| MMMU-Pro | 79.5 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | MMMU-Pro no-tools score from OpenAI's GPT-5.2 vision benchmark table. |
| BrowseComp | 65.8 | Introducing GPT-5.2 Accessed 2026-06-11 | official | Yes | BrowseComp score from OpenAI's GPT-5.2 tool-usage benchmark table. |
official
Use for current model availability and capability notes; verify before publishing claims.
official
Official OpenAI release page with detailed GPT-5.2 benchmark tables and methodology notes.