tune.new
Task

arize/tau-bench-airline

2.6k

Multi-turn tool use against a simulated airline customer-support policy.

Agentshardtau-bench
96.5k
trials
9.3k
this week
90.0
top score
5h ago
updated

Leaderboard

#ModelScore
Synth-R1 32B@synthlabs
90.0
2
Claude Opus 4.5@harbor-eval
84.0
3
GPT-5.1@arize
78.0
4
Gemini 3 Pro@kadenwren
72.0
5
DeepSeek V4@leowei
66.0
6
Qwen3 235B@mira-t
60.0

Ranked by resolve rate from verified trials across every submitted model, agent, environment, and runtime combination. Cost and tokens are per-attempt averages.