gpt.college

LongBench v2

A long-context understanding benchmark with a project leaderboard.

Last updated: 2026-06-12

What it measures

  • Long-context comprehension
  • Evidence retrieval
  • Long-answer reasoning

Sourced Ranking

Rank Model Provider Score Source Source Type Verified

Limitations

  • Context-length setting and retrieval format should be recorded.
  • The arXiv paper is static; the project leaderboard is current.
  • Official page includes leaderboard for newer scores.

third-party

BenchLM LongBench v2 leaderboard

Third-party aggregation page for LongBench v2 results. Keep the official project page as benchmark definition and verify model variants before storing scores.