Wrong answer Failures
See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Total Cost ↑.
Categories
291/291
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #62 | Qwen3.8 27B low | Qwen | 4 | 7.9 | ~$0.071 | 15/22 | 39.1s |
| #94 | GPT-5.6 Luna medium | OpenAI | 9 | 7.4 | $0.072 | 13/22 | 7.43s |
| #153 | Qwen3.5 Plus 2026-02-15 none | Qwen | 11 | 6.6 | $0.073 | 11/22 | 9.86s |
| #189 | Qwen3.5-Flash none | Qwen | 14 | 6.0 | $0.073 | 7/22 | 25.3s |
| #119 | DeepSeek V3.2 medium | DeepSeek | 5 | 7.0 | $0.074 | 11/22 | 68.4s |
| #209 | KAT-Coder-Air V2.5 high | Kwaipilot | 9 | 5.6 | $0.077 | 7/22 | 15.9s |
| #246 | Laguna S 2.1 low | Poolside | 15 | 5.0 | $0.082 | 3/22 | 85.3s |
| #217 | Qwen3.6 27B none | Qwen | 11 | 5.5 | $0.083 | 7/22 | 10.6s |
| #138 | Gemini 3 Flash Preview none | 8 | 6.8 | $0.085 | 13/22 | 2.82s | |
| #268 | Grok 4.20 Beta none | X AI | 10 | 4.4 | $0.087 | 6/18 | 1.19s |
| #197 | Seed 2.1 Turbo none | Bytedance Seed | 11 | 5.9 | $0.092 | 10/22 | 6.84s |
| #139 | Gemma 4 26B A4B medium | 2 | 6.8 | $0.092 | 15/22 | 104.9s | |
| #122 | Mercury 2 medium | Inception | 8 | 7.0 | $0.094 | 10/22 | 2.95s |
| #182 | Nemotron 3 Ultra none | NVIDIA | 11 | 6.1 | $0.095 | 8/22 | 3.88s |
| #196 | GPT-5.4 Mini none | OpenAI | 13 | 5.9 | $0.095 | 6/22 | 1.58s |