gpt.college

BrowseComp

A browsing benchmark for hard-to-find information and deep web research tasks.

Last updated: 2026-06-11

What it measures

  • Web browsing
  • Deep search
  • Evidence gathering

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 Gemini 3.1 Pro Google 85.9 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes
2 OpenAI GPT-5.4 OpenAI 82.7 Introducing GPT-5.4 Accessed 2026-06-11 official Yes
3 OpenAI GPT-5.2 Thinking OpenAI 65.8 Introducing GPT-5.2 Accessed 2026-06-11 official Yes
4 Gemini 3 Pro Google 59.2 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes

Limitations

  • Cite model-release score sources separately.
  • The arXiv paper is static; latest scores usually appear in model release pages.
  • OpenAI page and code are official entry points, but not a single dynamic leaderboard.

leaderboard

OpenAI BrowseComp page and simple-evals code

The arXiv paper is static; latest scores usually appear in model release pages. OpenAI page and code are official entry points, but not a single dynamic leaderboard.

third-party

BrowseComp dataset or code

Use for dataset, code, or implementation details; score freshness depends on the benchmark.