What it measures
- Factuality
- Short-answer knowledge
- Hallucination resistance
A factuality benchmark for short-answer factual knowledge questions.
% correct; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The arXiv paper is static; Kaggle is the operational source. Google DeepMind points to Kaggle for dataset, eval, and leaderboard entry.
paper
Use for benchmark definition, not necessarily for latest model scores.
official
Use for dataset, code, or implementation details; score freshness depends on the benchmark.
leaderboard
Related SimpleQA benchmark page from OpenAI. Do not merge SimpleQA and SimpleQA Verified rows without explicit labeling.