Trivia: Wrong answer
Trivia
Wrong answer
See which AI models are most likely to hit Wrong answer on Trivia, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
197/197
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #193 | Laguna XS 2.1 none | Poolside | 1 | 3.0 | $0.008 | 0/1 | 254ms |
| #171 | Qwen3.5-122B-A10B none | Qwen | 1 | 3.0 | $0.247 | 0/1 | 295ms |
| #237 | Granite 4.1 8B none | IBM Granite | 1 | 3.0 | $0.007 | 0/1 | 306ms |
| #198 | Mistral Small 4 none | Mistral | 1 | 3.0 | $0.022 | 0/1 | 397ms |
| #222 | Qwen3 Coder Next medium | Qwen | 1 | 3.0 | $0.034 | 0/1 | 399ms |
| #194 | Qwen3.6 35B A3B none | Qwen | 1 | 3.0 | $0.061 | 0/1 | 414ms |
| #192 | Inkling Small none | Thinkingmachines | 1 | 3.0 | $0.058 | 0/1 | 481ms |
| #152 | Qwen3.5-35B-A3B none | Qwen | 1 | 3.0 | $0.106 | 0/1 | 493ms |
| #224 | Mercury 2 none | Inception | 1 | 3.0 | $0.030 | 0/1 | 548ms |
| #150 | Qwen3.5-Flash none | Qwen | 1 | 3.0 | $0.073 | 0/1 | 588ms |
| #125 | Qwen3.5-27B none | Qwen | 1 | 3.0 | $0.058 | 0/1 | 599ms |
| #199 | Qwen3 Coder Next none | Qwen | 1 | 3.0 | $0.026 | 0/1 | 601ms |
| #149 | Qwen3.6 Flash none | Qwen | 1 | 3.0 | $0.062 | 0/1 | 649ms |
| #189 | GPT-5.6 Luna none | OpenAI | 1 | 3.0 | $0.015 | 0/1 | 651ms |
| #197 | Inkling none | Thinkingmachines | 1 | 3.0 | $0.147 | 0/1 | 670ms |