What it measures
- Web navigation
- Multi-step task completion
- Tool and browser control
A web-agent benchmark for completing tasks in realistic self-hosted web environments.
% success; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The paper is static; project and leaderboard sources are better for current results. Record self-hosted environment version when storing scores.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.
third-party
Third-party tracked leaderboard for WebArena agent setups. Use as a candidate score source only after verifying environment, agent scaffold, and submission details.