gpt.college

Benchmarks

Human-readable benchmark definitions with score direction, limitations, and source records.

score / higher is better

ARC-AGI-2

A hard abstract-reasoning benchmark from ARC Prize focused on novel task generalization.

% accuracy / higher is better

Humanity's Last Exam (HLE)

A broad expert-level benchmark intended to test difficult questions across many domains and formats.

% solved / higher is better

FrontierMath

An advanced mathematics benchmark curated around difficult research-style math problems.

% accuracy / higher is better

GPQA / GPQA Diamond

A graduate-level, Google-proof question-answering benchmark with a commonly used Diamond subset.

% accuracy / higher is better

MMLU-Pro

A harder MMLU-style benchmark for broad knowledge and reasoning.

% resolved / higher is better

SWE-bench Verified

A curated SWE-bench split for resolving real software issues with agentic coding systems.

% resolved / higher is better

SWE-bench Pro

A professional software-engineering agent benchmark with public and private leaderboard splits.

% success / higher is better

Terminal-Bench 2.x

A terminal-environment agent benchmark for command-line task completion.

medal or score / higher is better

MLE-bench

A machine-learning engineering benchmark based on Kaggle-style competition tasks.

% accuracy / higher is better

BrowseComp

A browsing benchmark for hard-to-find information and deep web research tasks.

% success / higher is better

tau-bench / tau2-bench

A tool-use benchmark family for business-process and customer-service style agent tasks.

% success / higher is better

WebArena

A web-agent benchmark for completing tasks in realistic self-hosted web environments.

% success / higher is better

VisualWebArena

A visual web-agent benchmark extending WebArena with image-grounded web tasks.

% accuracy / higher is better

LongBench v2

A long-context understanding benchmark with a project leaderboard.

% accuracy / higher is better

RULER

A synthetic long-context stress benchmark from NVIDIA.

score / higher is better

HELMET

A long-context benchmark suite focused on practical application-style tasks.

% accuracy / higher is better

MMMU-Pro

A harder multimodal academic benchmark derived from MMMU-style tasks.

% accuracy / higher is better

Video-MME

A video understanding benchmark for multimodal models.

% accuracy / higher is better

MathVista

A visual math reasoning benchmark for multimodal models.

% accuracy / higher is better

CharXiv

A chart and figure understanding benchmark for scientific visual reasoning.

% correct / higher is better

SimpleQA Verified

A factuality benchmark for short-answer factual knowledge questions.

% accuracy / higher is better

FRAMES

A RAG and factual multi-hop question answering benchmark.

score / higher is better

GraphWalks

A graph reasoning benchmark candidate that still needs a stable authoritative source before launch.

% success / higher is better

TerminalWorld-Verified

A terminal-agent benchmark candidate that needs source confirmation.