Strengths
- Strong SWE-Bench Pro result
- High GPQA Diamond result
- Broad multimodal and tool-use coverage
A frontier OpenAI model release with strong reported performance across coding, tool use, multimodal, and academic reasoning benchmarks.
| Benchmark | Score | Source | Source Type | Verified | Notes |
|---|---|---|---|---|---|
| SWE-bench Pro | 57.7 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes | SWE-Bench Pro public score reported by OpenAI for GPT-5.4. |
| GPQA / GPQA Diamond | 92.8 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes | GPQA Diamond score from OpenAI's GPT-5.4 academic benchmark table. |
| FrontierMath | 47.6 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes | FrontierMath Tier 1-3 score from OpenAI's GPT-5.4 academic benchmark table. |
| MMMU-Pro | 81.2 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes | MMMU-Pro no-tools score from OpenAI's GPT-5.4 computer-use and vision table. |
| BrowseComp | 82.7 | Introducing GPT-5.4 Accessed 2026-06-11 | official | Yes | BrowseComp score from OpenAI's GPT-5.4 tool-use table. |
| Humanity's Last Exam (HLE) | 36.47 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | Scale Labs HLE Text Only score for gpt-5.4-2026-03-05 (xhigh thinking), covering the 86% text-only HLE subset. |
official
Use for current model availability and capability notes; verify before publishing claims.
official
Official OpenAI release page with GPT-5.4 coding, tool-use, vision, and academic benchmark tables.