Coding
Signals for code generation, debugging, ML engineering, and terminal-based software agent work.
A source-backed radar view for comparing coding, reasoning, research, multimodal, and agentic benchmark signals. Missing dimensions are treated as missing evidence, not as hidden claims.
Each axis is a source-backed category rollup with visible coverage and confidence. Missing dimensions mean this registry does not yet have sourced evidence for that model and category.
Use-case lenses change which evidence matters most. These links go to category pages built from the same registry.
Signals for code generation, debugging, ML engineering, and terminal-based software agent work.
Signals for abstract reasoning, hard question answering, math, and multi-step problem solving.
Signals for synthesis, factual work, long answers, RAG, and source-heavy analysis.
Signals for models that need to understand images, charts, video, and text together.
Signals for tool use, browsing, GUI control, terminal operation, and multi-step autonomous workflows.
| Model | Coding | Reasoning | Research | Multimodal | Agentic | Evidence |
|---|---|---|---|---|---|---|
| OpenAI GPT-5.4 | 57.7 1 of 5 sourced rows limited | 59 3 of 6 sourced rows sufficient | 87.8 2 of 9 sourced rows limited | 58.8 2 of 7 sourced rows limited | 70.2 2 of 10 sourced rows limited | Verified source rows 5 of 5 axes with evidence |
| OpenAI GPT-5.2 Thinking | 67.8 2 of 5 sourced rows limited | 66.3 2 of 6 sourced rows limited | 79.1 2 of 9 sourced rows limited | 79.5 1 of 7 sourced rows limited | 67.1 3 of 10 sourced rows sufficient | Verified source rows 5 of 5 axes with evidence |
| Claude Opus 4.6 | 81.4 1 of 5 sourced rows limited | 36.2 1 of 6 sourced rows limited | No sourced row | 36.2 1 of 7 sourced rows limited | 81.4 1 of 10 sourced rows limited | Verified source rows 4 of 5 axes with evidence |
| Gemini 3.1 Pro | 67.8 3 of 5 sourced rows sufficient | 72.9 3 of 6 sourced rows sufficient | 90.1 2 of 9 sourced rows limited | 63.9 2 of 7 sourced rows limited | 72.3 4 of 10 sourced rows sufficient | Verified source rows 5 of 5 axes with evidence |
| DeepSeek-R1 | 49.2 1 of 5 sourced rows limited | 54.7 3 of 6 sourced rows sufficient | 79.3 3 of 9 sourced rows sufficient | 8.5 1 of 7 sourced rows limited | 49.2 1 of 10 sourced rows limited | Verified source rows 5 of 5 axes with evidence |
Capability rows use only `src/data/results.ts`. Demo names from the prototype, including Atlas, Forge, and Nova, are not used. Every visible score links back to a source record and carries a note about harness or methodology when relevant.