blog
Research notes and eval methodology.
How we measure frontier models and agents — and what we learn building the infrastructure to do it.
Methodology
How we score agentic coding without leaking the test suite
A look at the isolation model behind SWE-bench Verified and why reproducibility beats raw pass-rate.
Mar 4, 20268 min
Product
Trending vs. popular: ranking evals by how often they run
Why we surface 7-day run velocity alongside all-time volume, and what it tells you about a benchmark.
Feb 22, 20265 min
Research
Synth-R1 32B: a small model that tops frontier agent tasks
The post-training recipe and verifier stack behind our latest open-weight release.
Feb 9, 202612 min
Engineering
Reproducible runtimes: pinning vLLM, SGLang, and TGI for eval parity
Small runtime differences move scores by points. Here is how we pin everything.
Jan 28, 20267 min