What it measures
- Video comprehension
- Temporal reasoning
- Multimodal QA
A video understanding benchmark for multimodal models.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The arXiv paper is static; the official leaderboard is current. Record short, medium, long, with-subtitles, and no-subtitles splits.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.