# Model Agent Env Runtime Score Cost Trials
synth-scaffold harbor-sandbox synth-router 90.0 $0.22 4.1k 2 harbor-react docker/py3.12 vLLM 0.9 84.0 $0.31 3.6k 3 OpenHands docker/py3.12 SGLang 78.0 $0.40 3.0k 4 SWE-agent e2b TGI 3.0 72.0 $0.49 2.5k 5 aider modal vLLM 0.9 66.0 $0.58 2.0k 6 harbor-react harbor-sandbox SGLang 60.0 $0.67 1.4k Ranked by resolve rate from verified trials across every submitted model, agent, environment, and runtime combination. Cost and tokens are per-attempt averages.
Model Synth-R1 32B
Agent synth-scaffold
Environment harbor-sandbox
Runtime synth-router 90.0 Score
$0.22 Cost
17.0k Tokens
How ranking works Each model, agent, environment, and runtime combination needs at least 50 verified trials to appear. Ties break on cost, then tokens. Scores refresh as new trials land.