gpt.college

Multimodal Ranking

Signals for models that need to understand images, charts, video, and text together.

Main rankings require at least 3 of 7 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
OpenAI GPT-5.2 Thinking OpenAI 79.5 1/7 Introducing GPT-5.2 Accessed 2026-06-11 2026-06-11
Gemini 3.1 Pro Google 63.9 2/7 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12
Gemini 3 Pro Google 59.4 2/7 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12
OpenAI GPT-5.4 OpenAI 58.8 2/7 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12
Claude Opus 4.6 Anthropic 36.2 1/7 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12
DeepSeek-R1 DeepSeek 8.5 1/7 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 2026-06-12

% accuracy

Humanity's Last Exam (HLE)

A broad expert-level benchmark intended to test difficult questions across many domains and formats.

% success

VisualWebArena

A visual web-agent benchmark extending WebArena with image-grounded web tasks.

% accuracy

MMMU-Pro

A harder multimodal academic benchmark derived from MMMU-style tasks.

% accuracy

Video-MME

A video understanding benchmark for multimodal models.

% accuracy

MathVista

A visual math reasoning benchmark for multimodal models.

% accuracy

CharXiv

A chart and figure understanding benchmark for scientific visual reasoning.