gpt.college

OpenAI / active

OpenAI GPT-5.4

A frontier OpenAI model release with strong reported performance across coding, tool use, multimodal, and academic reasoning benchmarks.

codingreasoningmultimodalresearchagentic

Last updated: 2026-06-11

Strengths

  • Strong SWE-Bench Pro result
  • High GPQA Diamond result
  • Broad multimodal and tool-use coverage

Limitations

  • Scores are self-reported by OpenAI and should be compared with harness notes visible

Benchmark Results

Benchmark Score Source Source Type Verified Notes
SWE-bench Pro 57.7 Introducing GPT-5.4 Accessed 2026-06-11 official Yes SWE-Bench Pro public score reported by OpenAI for GPT-5.4.
GPQA / GPQA Diamond 92.8 Introducing GPT-5.4 Accessed 2026-06-11 official Yes GPQA Diamond score from OpenAI's GPT-5.4 academic benchmark table.
FrontierMath 47.6 Introducing GPT-5.4 Accessed 2026-06-11 official Yes FrontierMath Tier 1-3 score from OpenAI's GPT-5.4 academic benchmark table.
MMMU-Pro 81.2 Introducing GPT-5.4 Accessed 2026-06-11 official Yes MMMU-Pro no-tools score from OpenAI's GPT-5.4 computer-use and vision table.
BrowseComp 82.7 Introducing GPT-5.4 Accessed 2026-06-11 official Yes BrowseComp score from OpenAI's GPT-5.4 tool-use table.
Humanity's Last Exam (HLE) 36.47 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes Scale Labs HLE Text Only score for gpt-5.4-2026-03-05 (xhigh thinking), covering the 86% text-only HLE subset.

official

Introducing GPT-5.4

Official OpenAI release page with GPT-5.4 coding, tool-use, vision, and academic benchmark tables.