Auto-RL tuning
Tune a model on any task. Keep the weights.
Every tune job starts from verifiable trajectories and ends on a leaderboard. No reward hacking, no hidden eval, no vendor lock on the checkpoint.
01
Start from trajectories
Pick a task. The job pulls every published trajectory with its per-step reward, or your private ones.
02
Run GRPO or DPO
Rollouts are scored by the task's verifier in isolated environments. Checkpoints are evaluated on the leaderboard as they land.
03
Keep the weights
The output is a checkpoint in your tenant plus a scored entry on the task leaderboard. Baseline, delta, and cost are on the record.
Public jobsOpen dashboard
Tune jobs running right now
| Job | Base model · task | Method | Status | Progress | Reward | Δ score | Cost |
|---|---|---|---|---|---|---|---|
swe-verified-grpo-v4synthlabs/tune_7e1c02 | Qwen3 235BSWE-bench Verified | GRPO | running | 1,840/3,000 | +8.8 | $412.50 | |
airline-policy-dpoarize/tune_3a90f1 | Llama 4 405Bτ-bench Airline | DPO | completed | 1,200/1,200 | +17.8 | $268.00 | |
terminal-sft-rlkadenwren/tune_c04d88 | Synth-R1 32BTerminal-Bench | SFT → RL | completed | 2,400/2,400 | +17.4 | $190.20 | |
webarena-shop-grpomira-t/tune_51de07 | Gemini 3 ProWebArena Shopping | GRPO | failed | 310/1,500 | +0.7 | $61.70 |