gpt.college

Humanity's Last Exam (HLE)

A broad expert-level benchmark intended to test difficult questions across many domains and formats.

Last updated: 2026-06-12

What it measures

  • Expert knowledge
  • Multimodal question answering
  • Frontier-model saturation resistance

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 Gemini 3.1 Pro Google 47.31 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes
2 Gemini 3 Pro Google 37.72 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes
3 OpenAI GPT-5.4 OpenAI 36.47 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes
4 Claude Opus 4.6 Anthropic 36.24 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes
5 DeepSeek-R1 DeepSeek 8.54 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes

Limitations

  • Use the rolling dashboard for current scores and the paper for benchmark definition.
  • Separate text-only, multimodal, and rolling variants in result records.
  • The paper is static; use the Scale Labs text-only leaderboard for current text-only scores and the official hub for benchmark definition.
  • Text-only leaderboard covers the 86% text-only HLE subset; do not mix it with multimodal or HLE-Rolling scores without labeling.

leaderboard

Scale Labs HLE Text Only leaderboard

The paper is static; use the Scale Labs text-only leaderboard for current text-only scores and the official hub for benchmark definition. Text-only leaderboard covers the 86% text-only HLE subset; do not mix it with multimodal or HLE-Rolling scores without labeling.

official

Humanity's Last Exam official hub

Official benchmark hub for paper, dataset, and methodology links. Use the Scale Labs leaderboard for the text-only operational score table.