gpt.college

Rankings

Category pages group related benchmark signals from source-backed model-version records.

View Capability Map

5 main ranking / 1 limited evidence

Overall

A cautious editorial rollup of the selected benchmark set. It is not an absolute model quality score.

Score leader with ≥3 sourced rows: Gemini 3.1 Pro (73.6)

1 main ranking / 5 limited evidence

Coding

Signals for code generation, debugging, ML engineering, and terminal-based software agent work.

Score leader with ≥3 sourced rows: Gemini 3.1 Pro (67.8)

2 main ranking / 4 limited evidence

Agentic

Signals for tool use, browsing, GUI control, terminal operation, and multi-step autonomous workflows.

Score leader with ≥3 sourced rows: Gemini 3.1 Pro (72.3)

3 main ranking / 3 limited evidence

Reasoning

Signals for abstract reasoning, hard question answering, math, and multi-step problem solving.

Score leader with ≥3 sourced rows: Gemini 3.1 Pro (72.9)

0 main ranking / 2 limited evidence

Math

Signals for advanced mathematics, contest-style math, and visual math reasoning.

4 main ranking / 2 limited evidence

Knowledge

Signals for expert knowledge, broad academic coverage, and factual accuracy.

Score leader with ≥3 sourced rows: Gemini 3.1 Pro (74)

1 main ranking / 4 limited evidence

Research

Signals for synthesis, factual work, long answers, RAG, and source-heavy analysis.

Score leader with ≥3 sourced rows: DeepSeek-R1 (79.3)

0 main ranking / 6 limited evidence

Multimodal

Signals for models that need to understand images, charts, video, and text together.

0 main ranking / 0 limited evidence

Long Context

Signals for long-context comprehension, retrieval, and stress testing.

0 main ranking / 4 limited evidence

Browsing

Signals for web research, browser navigation, and web-agent task completion.

0 main ranking / 0 limited evidence

GUI / Computer Use

Signals for operating-system, graphical interface, and computer-use agents.

0 main ranking / 0 limited evidence

Video

Signals for video understanding and temporal multimodal reasoning.