gpt.college

DeepSeek / active

DeepSeek-R1

A reasoning model from DeepSeek with official benchmark results published on the DeepSeek-R1 model card.

codingreasoningmathresearch

Last updated: 2026-06-11

Strengths

  • Strong open reasoning model positioning
  • Reported GPQA, MMLU-Pro, FRAMES, and math benchmark results

Limitations

  • Hosted/API behavior and open model usage can differ by deployment

Benchmark Results

Benchmark Score Source Source Type Verified Notes
GPQA / GPQA Diamond 71.5 DeepSeek-R1 model card Accessed 2026-06-11 official Yes GPQA-Diamond Pass@1 score from the official DeepSeek-R1 model card.
MMLU-Pro 84 DeepSeek-R1 model card Accessed 2026-06-11 official Yes MMLU-Pro EM score from the official DeepSeek-R1 model card.
SWE-bench Verified 49.2 DeepSeek-R1 model card Accessed 2026-06-11 official Yes SWE Verified resolved score from the official DeepSeek-R1 model card.
FRAMES 82.5 DeepSeek-R1 model card Accessed 2026-06-11 official Yes FRAMES accuracy from the official DeepSeek-R1 model card.
Humanity's Last Exam (HLE) 8.54 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes Scale Labs HLE Text Only score for DeepSeek R1, covering the 86% text-only HLE subset.

official

DeepSeek-R1 model card

Official DeepSeek-R1 model card mirrored on Hugging Face with benchmark table for GPQA, MMLU-Pro, SWE Verified, FRAMES, and related tasks.