tune.new
Explore · Leaderboards

Leaderboards6 tasks

One leaderboard per task. Rankings are computed from verified trials, and every score links to the trajectories that produced it.

Codingharbor-eval/swe-bench-verified · 184.2k runs
Open task
#ModelScore
Synth-R1 32B@synthlabs
91.0
2
Claude Opus 4.5@harbor-eval
85.0
3
GPT-5.1@arize
79.0
4
Gemini 3 Pro@kadenwren
73.0
5
DeepSeek V4@leowei
67.0
6
Qwen3 235B@mira-t
61.0