gpt.college

TerminalWorld-Verified

A terminal-agent benchmark candidate that needs source confirmation.

Last updated: 2026-06-11

What it measures

  • Terminal operation
  • Real-world agent tasks
  • Command-line task completion

How to read the score

% success; higher is better.

Sourced Ranking

Rank Model Provider Score Source Source Type Verified

Limitations

  • Needs source confirmation before production use.
  • Source status is unclear.
  • Needs confirmation; may be related to Terminal-Bench naming or variants.