Trivia: Wrong answer
Trivia
Wrong answer
See which AI models are most likely to hit Wrong answer on Trivia, so you can spot weak points faster. Sort by: Response Time (avg) ↓.
Failure Reasons
197/197
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #245 | Step 3.5 Flash none | Stepfun | 1 | 3.0 | $0.020 | 0/1 | 114.1s |
| #132 | Ring-2.6-1T medium | Inclusionai | 1 | 3.0 | $0.103 | 0/1 | 113.9s |
| #46 | DeepSeek V4 Flash 0731 high | DeepSeek | 1 | 3.0 | $0.079 | 0/1 | 110.6s |
| #156 | Step 3.5 Flash medium | Stepfun | 1 | 3.0 | $0.108 | 0/1 | 108.4s |
| #75 | Qwen3.5 Plus 2026-02-15 medium | Qwen | 1 | 3.0 | $0.437 | 0/1 | 103.8s |
| #160 | Trinity Large Thinking medium | Arcee AI | 1 | 3.0 | $0.792 | 0/1 | 97.4s |
| #90 | Qwen3.5 Plus 2026-04-20 medium | Qwen | 1 | 3.0 | $0.317 | 0/1 | 92.6s |
| #51 | Qwen3.7 Plus medium | Qwen | 1 | 3.0 | $0.267 | 0/1 | 91.1s |
| #134 | Gemma 4 31B medium | 1 | 3.0 | $0.097 | 0/1 | 90.1s | |
| #76 | Qwen3.5-27B medium | Qwen | 1 | 3.0 | $0.981 | 0/1 | 85.1s |
| #96 | DeepSeek V3.2 medium | DeepSeek | 1 | 3.0 | $0.078 | 0/1 | 84.0s |
| #97 | Kimi K2.5 medium | Moonshot AI | 1 | 3.0 | $0.600 | 0/1 | 83.9s |
| #133 | Mimo V2 PRO medium | Xiaomi | 1 | 3.0 | $0.333 | 0/1 | 82.7s |
| #121 | Qwen3.6 27B medium | Qwen | 1 | 3.0 | $0.680 | 0/1 | 81.0s |
| #225 | MiniMax M2.5 medium | Minimax | 1 | 3.0 | $0.350 | 0/1 | 80.8s |