Combined: Wrong answer
Combined
Wrong answer
See which AI models are most likely to hit Wrong answer on Combined, so you can spot weak points faster. Sort by: Total Cost ↓.
Failure Reasons
64/64
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #171 | Mistral Small 4 none | Mistral | 2 | 3.0 | $0.022 | 0/2 | 7.44s |
| #163 | Mimo V2 Omni none | Xiaomi | 1 | 1.5 | $0.021 | 0/1 | 5.96s |
| #121 | Gemma 4 31B none | 1 | 3.8 | $0.021 | 0/2 | 30.0s | |
| #124 | Gemini 2.5 Flash none | 1 | 3.0 | $0.017 | 0/2 | 61.2s | |
| #168 | Ling-2.6-1T none | Inclusionai | 1 | 6.5 | $0.016 | 1/2 | 23.8s |
| #204 | Laguna Xs.2 medium | Poolside | 1 | 1.5 | $0.015 | 0/1 | 15.9s |
| #162 | Gemma 4 26B A4B none | 1 | 3.0 | $0.015 | 0/2 | 37.2s | |
| #180 | GPT-4o-mini none | OpenAI | 1 | 3.0 | $0.010 | 0/2 | 6.32s |
| #183 | Nemotron 3 Super none | NVIDIA | 2 | 3.0 | $0.008 | 0/2 | 18.2s |
| #189 | Trinity Large Preview none | Arcee AI | 1 | 1.5 | $0.008 | 0/1 | 8.91s |
| #209 | Grok 4.1 Fast none | X AI | 1 | 1.5 | $0.008 | 0/1 | 3.33s |
| #166 | Laguna XS 2.1 none | Poolside | 1 | 3.0 | $0.008 | 0/2 | 10.4s |
| #211 | Laguna Xs.2 none | Poolside | 1 | 1.5 | $0.004 | 0/1 | 2.01s |
| #205 | Hy3 preview none | Tencent | 1 | 1.5 | $0.003 | 0/1 | 35.8s |
| #152 | Owl Alpha medium | Openrouter | 1 | 1.5 | $0.000 | 0/1 | 10.0s |