gpt.college

MMLU-Pro

A harder MMLU-style benchmark for broad knowledge and reasoning.

Last updated: 2026-06-11

What it measures

  • Knowledge breadth
  • Reasoning under harder answer choices
  • Academic subject coverage

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 DeepSeek-R1 DeepSeek 84 DeepSeek-R1 model card Accessed 2026-06-11 official Yes

Limitations

  • Confirm prompt format and model variant before comparing scores.
  • The arXiv paper is static; the leaderboard and GitHub may update.
  • HF Space is the score entry point; GitHub is the implementation entry point.

third-party

MMLU-Pro dataset or code

Use for dataset, code, or implementation details; score freshness depends on the benchmark.