Humanity's Last Exam (HLE)
A broad expert-level benchmark intended to test difficult questions across many domains and formats.
Signals for models that need to understand images, charts, video, and text together.
| Rank | Model | Provider | Score | Coverage | Source | Source Type | Verified | Last Updated |
|---|
| Model | Provider | Score | Coverage | Freshest Source | Last Updated |
|---|---|---|---|---|---|
| OpenAI GPT-5.2 Thinking | OpenAI | 79.5 | 1/7 | Introducing GPT-5.2 Accessed 2026-06-11 | 2026-06-11 |
| Gemini 3.1 Pro | 63.9 | 2/7 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 | |
| Gemini 3 Pro | 59.4 | 2/7 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 | |
| OpenAI GPT-5.4 | OpenAI | 58.8 | 2/7 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 |
| Claude Opus 4.6 | Anthropic | 36.2 | 1/7 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 |
| DeepSeek-R1 | DeepSeek | 8.5 | 1/7 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | 2026-06-12 |
A broad expert-level benchmark intended to test difficult questions across many domains and formats.
A computer-use benchmark for operating-system and GUI agents.
A visual web-agent benchmark extending WebArena with image-grounded web tasks.
A harder multimodal academic benchmark derived from MMMU-style tasks.
A video understanding benchmark for multimodal models.
A visual math reasoning benchmark for multimodal models.
A chart and figure understanding benchmark for scientific visual reasoning.