AI BENCHY Category Failures
Domain specific: Wrong answer
Domain specific
Wrong answer
See which AI models are most likely to hit Wrong answer on Domain specific, so you can spot weak points faster. Sort by: Tests Correct ↓.
Failure Reasons
| Rank | Model | Company | Wrong answer Count | Category Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #135 | Kimi K2.5 none | Moonshot AI | 2 | 5.3 | 1/3 | 4.38s |
| #139 | DeepSeek V4 Flash none | DeepSeek | 2 | 5.3 | 1/3 | 19.7s |
| #140 | Qwen3 Coder Next none | Qwen | 2 | 5.3 | 1/3 | 962ms |
| #142 | Mistral Small 4 none | Mistral | 2 | 5.3 | 1/3 | 367ms |
| #146 | Laguna Xs.2 none | Poolside | 2 | 5.3 | 1/3 | 371ms |
| #150 | Qwen3 Coder Next medium | Qwen | 2 | 5.3 | 1/3 | 638ms |
| #151 | Trinity Large Preview none | Arcee AI | 2 | 5.3 | 1/3 | 877ms |
| #152 | MiMo-V2-Flash none | Xiaomi | 2 | 5.3 | 1/3 | 564ms |
| #155 | Mercury 2 none | Inception | 2 | 5.3 | 1/3 | 534ms |
| #157 | Grok 4.1 Fast none | X AI | 2 | 5.9 | 1/3 | 1.06s |
| #160 | LFM2-24B-A2B none | Liquid | 1 | 5.9 | 1/3 | 287ms |
| #14 | Qwen3.6 Max Preview medium | Qwen | 3 | 2.9 | 0/3 | 95.9s |
| #17 | GLM 5 medium | Z.ai | 2 | 3.5 | 0/3 | 0ms |
| #18 | Qwen3.7 Plus medium | Qwen | 3 | 3.6 | 0/3 | 45.3s |
| #23 | GLM 5 Turbo medium | Z.ai | 2 | 2.9 | 0/3 | 71.1s |