Strengths
- Strong official SWE-bench Verified result
- Agentic workflow positioning
- Long-context workflow emphasis
Anthropic's Opus 4.6 release, positioned around long-running agentic coding, computer-use, and professional workflows.
| Benchmark | Score | Source | Source Type | Verified | Notes |
|---|---|---|---|---|---|
| SWE-bench Verified | 81.42 | Introducing Claude Opus 4.6 Accessed 2026-06-11 | official | Yes | Anthropic reports this SWE-bench Verified score averaged over 25 trials with a prompt modification. |
| Humanity's Last Exam (HLE) | 36.24 | Scale Labs HLE Text Only leaderboard Accessed 2026-06-12 | leaderboard | Yes | Scale Labs HLE Text Only score for claude-opus-4-6-thinking-max, covering the 86% text-only HLE subset. |
official
Use for current model family and capability notes; verify before publishing claims.
official
Official Anthropic release page. Use benchmark notes carefully because some reported scores use repeated trials or prompt modifications.