BrowseComp
A browsing benchmark for hard-to-find information and deep web research tasks.
Signals for web research, browser navigation, and web-agent task completion.
| Rank | Model | Provider | Score | Coverage | Source | Source Type | Verified | Last Updated |
|---|
| Model | Provider | Score | Coverage | Freshest Source | Last Updated |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 85.9 | 1/3 | Gemini 3.1 Pro model page Accessed 2026-06-11 | 2026-06-11 | |
| OpenAI GPT-5.4 | OpenAI | 82.7 | 1/3 | Introducing GPT-5.4 Accessed 2026-06-11 | 2026-06-11 |
| OpenAI GPT-5.2 Thinking | OpenAI | 65.8 | 1/3 | Introducing GPT-5.2 Accessed 2026-06-11 | 2026-06-11 |
| Gemini 3 Pro | 59.2 | 1/3 | Gemini 3.1 Pro model page Accessed 2026-06-11 | 2026-06-11 |
A browsing benchmark for hard-to-find information and deep web research tasks.
A web-agent benchmark for completing tasks in realistic self-hosted web environments.
A visual web-agent benchmark extending WebArena with image-grounded web tasks.