Wrong answer Failures
See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Total Cost ↓.
Categories
291/291
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #278 | Ling 3.0 Tiny none | Inclusionai | 13 | 4.0 | $0.000 | 2/22 | 14.0s |
| #282 | Ling 3.0 Tiny high | Inclusionai | 6 | 3.8 | $0.000 | 4/22 | 75.7s |
| #283 | Ling 3.0 Tiny low | Inclusionai | 5 | 3.8 | $0.000 | 4/22 | 66.2s |
| #286 | Ling 3.0 Tiny medium | Inclusionai | 5 | 3.8 | $0.000 | 4/22 | 68.1s |
| #289 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 7 | 3.4 | $0.000 | 4/19 | 17.1s |
| #290 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 9 | 3.2 | $0.000 | 2/19 | 728ms |