Explore · Benchmarks
Benchmarks6
A benchmark groups tasks into a suite. Trending is ranked by trials in the last seven days; popular by all-time trials.
6 results
harbor-eval/swe-bench13.8k
The reference suite for autonomous software engineering on real repositories.
codingagentsverified
- 512.0k
- trials
- 38.4k
- this week
- 12
- tasks
synthlabs/math-frontier7.8k
Verifier-checked competition math from AIME through Olympiad problems.
mathverifiedreasoning
- 289.0k
- trials
- 31.2k
- this week
- 8
- tasks
arize/tau-bench5.8k
Tool-agent-user interaction under realistic, policy-constrained scenarios.
agentstool-usemulti-turn
- 214.0k
- trials
- 19.8k
- this week
- 6
- tasks
harbor-eval/gaia4.8k
General assistant tasks that resist memorization and reward real reasoning.
reasoningtoolsgeneral
- 178.0k
- trials
- 14.2k
- this week
- 9
- tasks
leowei/webarena3.3k
Realistic, reproducible web agents in self-hosted application sandboxes.
webagentssandbox
- 121.0k
- trials
- 8.6k
- this week
- 7
- tasks
kadenwren/terminal-bench2.4k
Shell-native tasks that measure real command-line problem solving.
terminalcodingsandbox
- 88.0k
- trials
- 5.1k
- this week
- 5
- tasks