Combined: Wrong answer
Combined
Wrong answer
See which AI models are most likely to hit Wrong answer on Combined, so you can spot weak points faster. Sort by: Tests Correct ↑.
Failure Reasons
64/64
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #201 | Elephant Alpha medium | Openrouter | 1 | 1.5 | $0.000 | 0/1 | 3.70s |
| #202 | Hunter Alpha none | OpenRouter | 1 | 1.5 | $0.000 | 0/1 | 15.2s |
| #204 | Laguna Xs.2 medium | Poolside | 1 | 1.5 | $0.015 | 0/1 | 15.9s |
| #205 | Hy3 preview none | Tencent | 1 | 1.5 | $0.003 | 0/1 | 35.8s |
| #206 | MiMo-V2-Flash none | Xiaomi | 1 | 1.5 | $0.025 | 0/1 | 2.87s |
| #209 | Grok 4.1 Fast none | X AI | 1 | 1.5 | $0.008 | 0/1 | 3.33s |
| #211 | Laguna Xs.2 none | Poolside | 1 | 1.5 | $0.004 | 0/1 | 2.01s |
| #23 | Grok 4.5 low | X AI | 1 | 6.5 | $0.935 | 1/2 | 12.8s |
| #56 | Kimi K2.7 Code medium | Moonshot AI | 1 | 7.3 | $0.740 | 1/2 | 66.0s |
| #63 | Qwen3.7 Max none | Qwen | 1 | 6.5 | $0.197 | 1/2 | 37.2s |
| #87 | GPT-5.6 Sol none | OpenAI | 1 | 6.5 | $0.524 | 1/2 | 8.37s |
| #91 | GPT-5.5 none | OpenAI | 1 | 6.5 | $0.544 | 1/2 | 8.90s |
| #103 | Qwen3.6 Max Preview none | Qwen | 1 | 6.5 | $0.231 | 1/2 | 61.6s |
| #109 | Qwen3.5-27B none | Qwen | 1 | 6.4 | $0.090 | 1/2 | 39.4s |
| #113 | Qwen3.5 Plus 2026-02-15 none | Qwen | 1 | 6.5 | $0.073 | 1/2 | 64.8s |