gpt.college

Anthropic / active

Claude Opus 4.6

Anthropic's Opus 4.6 release, positioned around long-running agentic coding, computer-use, and professional workflows.

codingreasoningwritingresearchagentic

Last updated: 2026-06-11

Strengths

  • Strong official SWE-bench Verified result
  • Agentic workflow positioning
  • Long-context workflow emphasis

Limitations

  • Some benchmark notes include repeated trials or prompt modifications; compare harnesses carefully

Benchmark Results

Benchmark Score Source Source Type Verified Notes
SWE-bench Verified 81.42 Introducing Claude Opus 4.6 Accessed 2026-06-11 official Yes Anthropic reports this SWE-bench Verified score averaged over 25 trials with a prompt modification.
Humanity's Last Exam (HLE) 36.24 Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 leaderboard Yes Scale Labs HLE Text Only score for claude-opus-4-6-thinking-max, covering the 86% text-only HLE subset.

official

Claude models overview

Use for current model family and capability notes; verify before publishing claims.

official

Introducing Claude Opus 4.6

Official Anthropic release page. Use benchmark notes carefully because some reported scores use repeated trials or prompt modifications.