What it measures
- Visual math
- Diagram reasoning
- Quantitative multimodal QA
A visual math reasoning benchmark for multimodal models.
% accuracy; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The paper is static; the official leaderboard is current. Track testmini and test score definitions separately.
paper
Use for benchmark definition, not necessarily for latest model scores.
official
Use for dataset, code, or implementation details; score freshness depends on the benchmark.
leaderboard
Kaggle benchmark page can be used as an operational leaderboard source. Preserve split labels such as testmini and test when adding scores.