What it measures
- Visual web navigation
- Multimodal grounding
- Browser task completion
A visual web-agent benchmark extending WebArena with image-grounded web tasks.
% success; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The paper is static; project and GitHub sources are more current. WebArena-x is the shared entry point; VisualWebArena adds visual grounding.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.