What it measures
- Tool use
- Business workflows
- Multi-turn agent reliability
A tool-use benchmark family for business-process and customer-service style agent tasks.
% success; higher is better.
| Rank | Model | Provider | Score | Source | Source Type | Verified |
|---|
leaderboard
The paper is static; the project page is better for variants. Track tau-bench, tau2, tau3, banking, and voice variants separately.
paper
Use for benchmark definition, not necessarily for latest model scores.
third-party
Use for dataset, code, or implementation details; score freshness depends on the benchmark.
third-party
Third-party tracked leaderboard. Verify harness, domain, and task variant before using any score as a primary claim.