Wrong answer Failures
See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Total Cost ↓.
Categories
291/291
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #104 | Qwen3.7 Plus none | Qwen | 9 | 7.3 | $0.106 | 12/22 | 12.1s |
| #162 | MiMo-V2.5 medium | Xiaomi | 5 | 6.4 | $0.104 | 11/22 | 46.3s |
| #13 | Gemini 3.7 Flash low | 2 | 9.2 | $0.104 | 20/22 | 3.85s | |
| #164 | Ring-2.6-1T medium | Inclusionai | 6 | 6.4 | $0.102 | 11/22 | 68.8s |
| #168 | Gemma 4 31B medium | 3 | 6.2 | $0.100 | 13/22 | 78.6s | |
| #239 | Mistral Small 4 medium | Mistral | 13 | 5.1 | $0.097 | 5/22 | 10.8s |
| #196 | GPT-5.4 Mini none | OpenAI | 13 | 5.9 | $0.095 | 6/22 | 1.58s |
| #182 | Nemotron 3 Ultra none | NVIDIA | 11 | 6.1 | $0.095 | 8/22 | 3.88s |
| #122 | Mercury 2 medium | Inception | 8 | 7.0 | $0.094 | 10/22 | 2.95s |
| #139 | Gemma 4 26B A4B medium | 2 | 6.8 | $0.092 | 15/22 | 104.9s | |
| #197 | Seed 2.1 Turbo none | Bytedance Seed | 11 | 5.9 | $0.092 | 10/22 | 6.84s |
| #268 | Grok 4.20 Beta none | X AI | 10 | 4.4 | $0.087 | 6/18 | 1.19s |
| #138 | Gemini 3 Flash Preview none | 8 | 6.8 | $0.085 | 13/22 | 2.82s | |
| #217 | Qwen3.6 27B none | Qwen | 11 | 5.5 | $0.083 | 7/22 | 10.6s |
| #246 | Laguna S 2.1 low | Poolside | 15 | 5.0 | $0.082 | 3/22 | 85.3s |