What it measures
- Synthetic long-context pressure
- Needle retrieval
- Context scaling
A synthetic long-context stress benchmark from NVIDIA.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
unknown
The arXiv paper is static and there is no single official dynamic leaderboard. Best for analysis and self-run evaluation; model-release scores require separate sources.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.