gpt.college

Agentic Ranking

Signals for tool use, browsing, GUI control, terminal operation, and multi-step autonomous workflows.

Main rankings require at least 3 of 10 included sourced benchmark rows. Within the main ranking, rows are sorted by average normalized score; coverage is shown separately and used only as a tie-breaker. Missing benchmark data is not treated as zero; models below the threshold appear under limited evidence.
Rank Model Provider Score Coverage Source Source Type Verified Last Updated
1 Gemini 3.1 Pro Google 72.3 4/10 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes 2026-06-11
2 OpenAI GPT-5.2 Thinking OpenAI 67.1 3/10 Introducing GPT-5.2 Accessed 2026-06-11 official Yes 2026-06-11

Limited Evidence

These models have fewer than 3 sourced rows in this category.

Model Provider Score Coverage Freshest Source Last Updated
Claude Opus 4.6 Anthropic 81.4 1/10 Introducing Claude Opus 4.6 Accessed 2026-06-11 2026-06-11
OpenAI GPT-5.4 OpenAI 70.2 2/10 Introducing GPT-5.4 Accessed 2026-06-11 2026-06-11
Gemini 3 Pro Google 67.7 2/10 Gemini 3.1 Pro model page Accessed 2026-06-11 2026-06-11
DeepSeek-R1 DeepSeek 49.2 1/10 DeepSeek-R1 model card Accessed 2026-06-11 2026-06-11

% resolved

SWE-bench Verified

A curated SWE-bench split for resolving real software issues with agentic coding systems.

% resolved

SWE-bench Pro

A professional software-engineering agent benchmark with public and private leaderboard splits.

% success

Terminal-Bench 2.x

A terminal-environment agent benchmark for command-line task completion.

medal or score

MLE-bench

A machine-learning engineering benchmark based on Kaggle-style competition tasks.

% accuracy

BrowseComp

A browsing benchmark for hard-to-find information and deep web research tasks.

% success

tau-bench / tau2-bench

A tool-use benchmark family for business-process and customer-service style agent tasks.

% success

WebArena

A web-agent benchmark for completing tasks in realistic self-hosted web environments.

% success

VisualWebArena

A visual web-agent benchmark extending WebArena with image-grounded web tasks.