Explore · Models
Models8
Aggregate resolve rate across every task a model has been scored on. The more it is evaluated, the more its ranking is trusted.
8 results
synthlabs/synth-r1-32b658
Synth-R1 32B by SynthLabs. Scored on 6 tasks across the task suite.
- 90.2
- avg score
- 6
- tasks
- 24.4k
- trials
anthropic/claude-opus-4.5571
Claude Opus 4.5 by Anthropic. Scored on 6 tasks across the task suite.
- 84.2
- avg score
- 6
- tasks
- 21.1k
- trials
openai/gpt-5.1483
GPT-5.1 by OpenAI. Scored on 6 tasks across the task suite.
- 78.2
- avg score
- 6
- tasks
- 17.9k
- trials
google/gemini-3-pro396
Gemini 3 Pro by Google. Scored on 6 tasks across the task suite.
- 72.2
- avg score
- 6
- tasks
- 14.6k
- trials
deepseek/deepseek-v4308
DeepSeek V4 by DeepSeek. Scored on 6 tasks across the task suite.
- 66.2
- avg score
- 6
- tasks
- 11.4k
- trials
alibaba/qwen3-235b221
Qwen3 235B by Alibaba. Scored on 6 tasks across the task suite.
- 60.2
- avg score
- 6
- tasks
- 8.2k
- trials
meta/llama-4-405b108
Llama 4 405B by Meta. Scored on 0 tasks across the task suite.
- 0.0
- avg score
- 0
- tasks
- 0
- trials
mistral/mistral-large-3108
Mistral Large 3 by Mistral. Scored on 0 tasks across the task suite.
- 0.0
- avg score
- 0
- tasks
- 0
- trials