gpt.college

Model capability comparison

See model strengths as a shape, not a slogan.

A source-backed radar view for comparing coding, reasoning, research, multimodal, and agentic benchmark signals. Missing dimensions are treated as missing evidence, not as hidden claims.

Coding Reasoning Research Multimodal Agentic

Radar is useful only when the evidence is visible.

Each axis is a source-backed category rollup with visible coverage and confidence. Missing dimensions mean this registry does not yet have sourced evidence for that model and category.

OpenAI GPT-5.4 6 sourced benchmark rows
OpenAI GPT-5.2 Thinking 6 sourced benchmark rows
Claude Opus 4.6 2 sourced benchmark rows
Gemini 3.1 Pro 8 sourced benchmark rows
DeepSeek-R1 5 sourced benchmark rows

Task Lenses

Use-case lenses change which evidence matters most. These links go to category pages built from the same registry.

5 benchmark definitions

Coding

Signals for code generation, debugging, ML engineering, and terminal-based software agent work.

6 benchmark definitions

Reasoning

Signals for abstract reasoning, hard question answering, math, and multi-step problem solving.

9 benchmark definitions

Research

Signals for synthesis, factual work, long answers, RAG, and source-heavy analysis.

7 benchmark definitions

Multimodal

Signals for models that need to understand images, charts, video, and text together.

10 benchmark definitions

Agentic

Signals for tool use, browsing, GUI control, terminal operation, and multi-step autonomous workflows.

Capability Matrix

Model CodingReasoningResearchMultimodalAgentic Evidence
OpenAI GPT-5.4 57.7 1 of 5 sourced rows limited 59 3 of 6 sourced rows sufficient 87.8 2 of 9 sourced rows limited 58.8 2 of 7 sourced rows limited 70.2 2 of 10 sourced rows limited Verified source rows 5 of 5 axes with evidence
OpenAI GPT-5.2 Thinking 67.8 2 of 5 sourced rows limited 66.3 2 of 6 sourced rows limited 79.1 2 of 9 sourced rows limited 79.5 1 of 7 sourced rows limited 67.1 3 of 10 sourced rows sufficient Verified source rows 5 of 5 axes with evidence
Claude Opus 4.6 81.4 1 of 5 sourced rows limited 36.2 1 of 6 sourced rows limited No sourced row 36.2 1 of 7 sourced rows limited 81.4 1 of 10 sourced rows limited Verified source rows 4 of 5 axes with evidence
Gemini 3.1 Pro 67.8 3 of 5 sourced rows sufficient 72.9 3 of 6 sourced rows sufficient 90.1 2 of 9 sourced rows limited 63.9 2 of 7 sourced rows limited 72.3 4 of 10 sourced rows sufficient Verified source rows 5 of 5 axes with evidence
DeepSeek-R1 49.2 1 of 5 sourced rows limited 54.7 3 of 6 sourced rows sufficient 79.3 3 of 9 sourced rows sufficient 8.5 1 of 7 sourced rows limited 49.2 1 of 10 sourced rows limited Verified source rows 5 of 5 axes with evidence

Evidence Rules

Capability rows use only `src/data/results.ts`. Demo names from the prototype, including Atlas, Forge, and Nova, are not used. Every visible score links back to a source record and carries a note about harness or methodology when relevant.