Wrong answer Failures
See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Score ↓.
Categories
313/313
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #184 | Gemma 4 31B medium | 3 | 6.2 | $0.100 | 13/22 | 78.6s | |
| #185 | GPT-5.6 Luna low | OpenAI | 10 | 6.2 | $0.051 | 10/22 | 5.20s |
| #186 | Claude Sonnet 5 none | Anthropic | 8 | 6.1 | $0.548 | 7/22 | 6.05s |
| #187 | Gemini 3.1 Flash Lite minimal | 8 | 6.1 | $0.047 | 10/22 | 1.85s | |
| #188 | Qwen3.5-35B-A3B medium | Qwen | 2 | 6.1 | $0.675 | 11/22 | 110.8s |
| #189 | gpt-oss-120b medium | OpenAI | 9 | 6.1 | $0.020 | 9/22 | 20.8s |
| #190 | Qwen3.7 Flash none | Qwen | 13 | 6.1 | $0.019 | 7/22 | 10.1s |
| #191 | Gemini 3.1 Flash Lite none | 11 | 6.1 | $0.046 | 9/22 | 1.75s | |
| #192 | Qwen3.5-Flash medium | Qwen | 5 | 6.1 | $0.142 | 11/22 | 84.4s |
| #193 | Inkling low | Thinkingmachines | 8 | 6.1 | $0.187 | 10/22 | 5.11s |
| #194 | Trinity Large Thinking medium | Arcee AI | 8 | 6.1 | $0.756 | 8/22 | 85.4s |
| #195 | Qwen3.6 Flash none | Qwen | 12 | 6.1 | $0.062 | 7/22 | 3.73s |
| #196 | Gemma 4 31B none | 10 | 6.1 | $0.016 | 9/22 | 5.41s | |
| #197 | Qwen3.5 Plus 2026-04-20 none | Qwen | 12 | 6.1 | $0.122 | 8/22 | 13.1s |
| #198 | Nemotron 3 Ultra none | NVIDIA | 11 | 6.1 | $0.093 | 8/22 | 3.88s |