What it measures
- Terminal operation
- Tool use
- Agentic task completion
A terminal-environment agent benchmark for command-line task completion.
% success; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 68.5 | Gemini 3.1 Pro model page Accessed 2026-06-11 | official | Yes |
leaderboard
The paper is static; tbench.ai reflects benchmark versions. Record 2.0, 2.1, 3.0, harness, and agent scaffold separately.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.