Overall
A cautious editorial rollup of the selected benchmark set. It is not an absolute model quality score.
Score leader with ≥3 sourced rows: Gemini 3.1 Pro (73.6)
Category pages group related benchmark signals from source-backed model-version records.
A cautious editorial rollup of the selected benchmark set. It is not an absolute model quality score.
Score leader with ≥3 sourced rows: Gemini 3.1 Pro (73.6)
Signals for code generation, debugging, ML engineering, and terminal-based software agent work.
Score leader with ≥3 sourced rows: Gemini 3.1 Pro (67.8)
Signals for tool use, browsing, GUI control, terminal operation, and multi-step autonomous workflows.
Score leader with ≥3 sourced rows: Gemini 3.1 Pro (72.3)
Signals for abstract reasoning, hard question answering, math, and multi-step problem solving.
Score leader with ≥3 sourced rows: Gemini 3.1 Pro (72.9)
Signals for advanced mathematics, contest-style math, and visual math reasoning.
Signals for expert knowledge, broad academic coverage, and factual accuracy.
Score leader with ≥3 sourced rows: Gemini 3.1 Pro (74)
Signals for synthesis, factual work, long answers, RAG, and source-heavy analysis.
Score leader with ≥3 sourced rows: DeepSeek-R1 (79.3)
Signals for models that need to understand images, charts, video, and text together.
Signals for long-context comprehension, retrieval, and stress testing.
Signals for web research, browser navigation, and web-agent task completion.
Signals for operating-system, graphical interface, and computer-use agents.
Signals for video understanding and temporal multimodal reasoning.