What it measures
- Terminal operation
- Real-world agent tasks
- Command-line task completion
A terminal-agent benchmark candidate that needs source confirmation.
% success; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
unknown
Source status is unclear. Needs confirmation; may be related to Terminal-Bench naming or variants.