What it measures
- Long-context applications
- Retrieval and reasoning
- Task-level robustness
A long-context benchmark suite focused on practical application-style tasks.
score; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The arXiv paper is static; the project page and spreadsheet may update. Official page includes overall results and spreadsheet links.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.