AI BENCHY Category Failures
Instructions following: Wrong answer
Instructions following
Wrong answer
See which AI models are most likely to hit Wrong answer on Instructions following, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
| Rank | Model | Company | Wrong answer Count | Category Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #93 | GLM 4.7 Flash medium | Z.ai | 1 | 6.2 | 1/2 | 2.97s |
| #36 | GPT-5.3 Chat none | OpenAI | 1 | 8.3 | 1/2 | 3.29s |
| #55 | MiMo-V2-Omni none | Xiaomi | 1 | 6.5 | 1/2 | 4.18s |
| #28 | GPT-5.2 Chat none | OpenAI | 1 | 7.5 | 1/2 | 5.46s |
| #92 | Qwen3 Coder Next medium | Qwen | 1 | 4.8 | 0/2 | 7.34s |
| #33 | GLM 5.1 medium | Z.ai | 1 | 6.4 | 1/2 | 7.47s |
| #87 | Qwen3 Coder Next none | Qwen | 2 | 4.8 | 0/2 | 7.71s |
| #59 | Qwen3.5-Flash none | Qwen | 1 | 6.3 | 1/2 | 8.81s |
| #80 | MiniMax M2.7 medium | Minimax | 1 | 3.7 | 0/2 | 12.6s |