Wrong answer Failures
See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Score ↓.
Categories
313/313
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #154 | Gemma 4 26B A4B medium | 2 | 6.8 | $0.092 | 15/22 | 104.9s | |
| #155 | GLM 5.2 none | Z.ai | 7 | 6.7 | $0.153 | 13/22 | 9.65s |
| #156 | KAT-Coder-Pro V2.5 none | Kwaipilot | 10 | 6.7 | $0.487 | 11/22 | 26.0s |
| #157 | LongCat 2.0 low | Meituan | 7 | 6.7 | $0.412 | 10/22 | 113.6s |
| #158 | Gemini 3.5 Flash minimal | 6 | 6.7 | $0.300 | 13/22 | 2.65s | |
| #159 | GLM 5V Turbo medium | Z.ai | 7 | 6.7 | $0.457 | 11/21 | 23.1s |
| #160 | LongCat 2.0 high | Meituan | 7 | 6.7 | $0.492 | 9/22 | 153.3s |
| #161 | Claude Opus 4.7 none | Anthropic | 3 | 6.6 | $0.505 | 16/19 | 3.02s |
| #162 | Gemini 3.5 Flash Lite low | 8 | 6.6 | $0.138 | 12/22 | 2.56s | |
| #163 | Qwen3.5-27B none | Qwen | 11 | 6.6 | $0.058 | 9/22 | 4.77s |
| #164 | Gemini 3.1 Flash Lite Preview low | 6 | 6.6 | $0.646 | 14/22 | 16.7s | |
| #165 | Inkling Small medium | Thinkingmachines | 6 | 6.6 | $0.113 | 10/22 | 6.30s |
| #166 | Gemini 3.1 Flash Lite low | 8 | 6.6 | $0.621 | 13/22 | 16.2s | |
| #167 | Qwen3.6 35B A3B medium | Qwen | 5 | 6.6 | $0.672 | 12/22 | 58.5s |
| #168 | Mercury 2.5 Preview high | Inception | 6 | 6.6 | $0.030 | 12/22 | 4.01s |