gpt.college

Browsing Ranking

Signals for web research, browser navigation, and web-agent task completion.

Main rankings require at least 3 of 3 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
Gemini 3.1 Pro Google 85.9 1/3 Gemini 3.1 Pro model page Accessed 2026-06-11 2026-06-11
OpenAI GPT-5.4 OpenAI 82.7 1/3 Introducing GPT-5.4 Accessed 2026-06-11 2026-06-11
OpenAI GPT-5.2 Thinking OpenAI 65.8 1/3 Introducing GPT-5.2 Accessed 2026-06-11 2026-06-11
Gemini 3 Pro Google 59.2 1/3 Gemini 3.1 Pro model page Accessed 2026-06-11 2026-06-11

% accuracy

BrowseComp

A browsing benchmark for hard-to-find information and deep web research tasks.

% success

WebArena

A web-agent benchmark for completing tasks in realistic self-hosted web environments.

% success

VisualWebArena

A visual web-agent benchmark extending WebArena with image-grounded web tasks.