gpt.college

Video-MME

A video understanding benchmark for multimodal models.

Last updated: 2026-06-11

What it measures

  • Video comprehension
  • Temporal reasoning
  • Multimodal QA

Sourced Ranking

Rank Model Provider Score Source Source Type Verified

Limitations

  • Video length and subtitle settings must be explicit.
  • The arXiv paper is static; the official leaderboard is current.
  • Record short, medium, long, with-subtitles, and no-subtitles splits.

leaderboard

Video-MME official leaderboard

The arXiv paper is static; the official leaderboard is current. Record short, medium, long, with-subtitles, and no-subtitles splits.

third-party

Video-MME dataset or code

Use for dataset, code, or implementation details; score freshness depends on the benchmark.