gpt.college

HELMET

A long-context benchmark suite focused on practical application-style tasks.

Last updated: 2026-06-11

What it measures

  • Long-context applications
  • Retrieval and reasoning
  • Task-level robustness

Sourced Ranking

Rank Model Provider Score Source Source Type Verified

Limitations

  • Track subtask mix and spreadsheet version.
  • The arXiv paper is static; the project page and spreadsheet may update.
  • Official page includes overall results and spreadsheet links.

third-party

HELMET dataset or code

Use for dataset, code, or implementation details; score freshness depends on the benchmark.