tune.new
Task

harbor-eval/gaia-level-3

2.0k

Long-horizon reasoning with web, file, and code tools; no shortcuts.

Reasoningfrontiergaia
74.3k
trials
8.1k
this week
89.0
top score
1d ago
updated

Leaderboard

#ModelScore
Synth-R1 32B@synthlabs
89.0
2
Claude Opus 4.5@harbor-eval
83.0
3
GPT-5.1@arize
77.0
4
Gemini 3 Pro@kadenwren
71.0
5
DeepSeek V4@leowei
65.0
6
Qwen3 235B@mira-t
59.0

Ranked by resolve rate from verified trials across every submitted model, agent, environment, and runtime combination. Cost and tokens are per-attempt averages.