gpt.college

Terminal-Bench 2.x

A terminal-environment agent benchmark for command-line task completion.

Last updated: 2026-06-11

What it measures

  • Terminal operation
  • Tool use
  • Agentic task completion

Sourced Ranking

Rank Model Provider Score Source Source Type Verified
1 Gemini 3.1 Pro Google 68.5 Gemini 3.1 Pro model page Accessed 2026-06-11 official Yes

Limitations

  • Version and harness differences should not be compared as one score.
  • The paper is static; tbench.ai reflects benchmark versions.
  • Record 2.0, 2.1, 3.0, harness, and agent scaffold separately.