AI BENCHY Category Failures
Puzzle Solving: Wrong answer
Puzzle Solving
Wrong answer
See which AI models are most likely to hit Wrong answer on Puzzle Solving, so you can spot weak points faster. Sort by: Tests Correct ↑.
Failure Reasons
| Rank | Model | Company | Wrong answer Count | Category Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #163 | Granite 4.1 8B none | IBM Granite | 2 | 3.2 | 0/3 | 608ms |
| #22 | Step 3.7 Flash medium | Stepfun | 2 | 5.7 | 1/3 | 6.19s |
| #38 | Grok 4.3 medium | X AI | 1 | 5.9 | 1/3 | 22.5s |
| #41 | Nemotron 3 Ultra 550b A55b medium | NVIDIA | 2 | 5.5 | 1/3 | 3.54s |
| #43 | MiMo-V2.5-Pro medium | Xiaomi | 1 | 6.7 | 1/3 | 5.31s |
| #54 | GPT-5 Mini medium | OpenAI | 1 | 5.6 | 1/3 | 15.2s |
| #57 | Step 3.7 Flash low | Stepfun | 2 | 5.5 | 1/3 | 1.84s |
| #60 | Kimi K2.6 medium | Moonshot AI | 1 | 6.0 | 1/3 | 25.1s |
| #62 | Step 3.5 Flash medium | Stepfun | 1 | 5.3 | 1/3 | 7.22s |
| #71 | Step 3.7 Flash high | Stepfun | 2 | 5.3 | 1/3 | 10.2s |
| #72 | DeepSeek V3.2 medium | DeepSeek | 1 | 7.0 | 1/3 | 37.7s |
| #75 | Ring-2.6-1T medium | Inclusionai | 1 | 5.9 | 1/3 | 20.7s |
| #76 | Kimi K2.5 medium | Moonshot AI | 1 | 5.3 | 1/3 | 43.2s |
| #79 | Hunter Alpha medium | OpenRouter | 1 | 6.1 | 1/3 | 5.35s |
| #80 | Mimo V2 Omni medium | Xiaomi | 1 | 5.9 | 1/3 | 2.38s |