What it measures
- Retrieval-augmented generation
- Multi-hop factual QA
- Evidence synthesis
A RAG and factual multi-hop question answering benchmark.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | DeepSeek-R1 | DeepSeek | 82.5 | DeepSeek-R1 model card Accessed 2026-06-11 | official | Yes |
unknown
The arXiv paper is static; Hugging Face is the dataset source, not a live score source. No official dynamic leaderboard was selected; newer scores need self-runs or model reports.
paper
Use for benchmark definition, not necessarily for latest model scores.
official
Use for dataset, code, or implementation details; score freshness depends on the benchmark.
third-party
Third-party leaderboard with limited model coverage. Use for discovery and verify against primary model reports before strong claims.