Strengths
- Strong open reasoning model positioning
- Reported GPQA, MMLU-Pro, FRAMES, and math benchmark results
A reasoning model from DeepSeek with official benchmark results published on the DeepSeek-R1 model card.
| Benchmark | Score | Source | Source Type | Verified | Notes |
|---|---|---|---|---|---|
| GPQA / GPQA Diamond | 71.5 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes | GPQA-Diamond Pass@1 score from the official DeepSeek-R1 model card. |
| MMLU-Pro | 84 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes | MMLU-Pro EM score from the official DeepSeek-R1 model card. |
| SWE-bench Verified | 49.2 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes | SWE Verified resolved score from the official DeepSeek-R1 model card. |
| FRAMES | 82.5 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes | FRAMES accuracy from the official DeepSeek-R1 model card. |
| Humanity's Last Exam (HLE) | 8.54 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | Scale Labs HLE Text Only score for DeepSeek R1, covering the 86% text-only HLE subset. |
official
Use for model availability and API behavior; verify before publishing claims.
official
Official DeepSeek-R1 model card mirrored on Hugging Face with benchmark table for GPQA, MMLU-Pro, SWE Verified, FRAMES, and related tasks.