What it measures
- Long-context comprehension
- Evidence retrieval
- Long-answer reasoning
A long-context understanding benchmark with a project leaderboard.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The arXiv paper is static; the project leaderboard is current. Official page includes leaderboard for newer scores.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.
third-party
Third-party aggregation page for LongBench v2 results. Keep the official project page as benchmark definition and verify model variants before storing scores.