AI BENCHY Category Failures
Puzzle Solving: Wrong answer
Puzzle Solving
Wrong answer
See which AI models are most likely to hit Wrong answer on Puzzle Solving, so you can spot weak points faster. Sort by: Response Time (avg) ↓.
Failure Reasons
| Rank | Model | Company | Wrong answer Count | Category Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #24 | GPT-5.2 Chat none | OpenAI | 1 | 7.7 | 2/3 | 4.10s |
| #135 | Kimi K2.5 none | Moonshot AI | 3 | 3.0 | 0/3 | 4.04s |
| #64 | MiMo-V2-Flash medium | Xiaomi | 1 | 7.7 | 2/3 | 3.87s |
| #70 | GPT-5.4 Nano medium | OpenAI | 2 | 4.1 | 0/3 | 3.79s |
| #116 | Hunter Alpha none | OpenRouter | 1 | 5.8 | 1/3 | 3.71s |
| #41 | Nemotron 3 Ultra 550b A55b medium | NVIDIA | 2 | 5.5 | 1/3 | 3.54s |
| #111 | Owl Alpha medium | Openrouter | 1 | 5.3 | 1/3 | 3.40s |
| #28 | Gemini 2.5 Flash medium | 1 | 7.7 | 2/3 | 3.18s | |
| #105 | Nemotron 3 Super medium | NVIDIA | 2 | 3.0 | 0/3 | 3.15s |
| #110 | Seed-2.0-Lite none | Bytedance Seed | 2 | 5.3 | 1/3 | 2.78s |
| #95 | Qwen3.5 Plus 2026-02-15 none | Qwen | 1 | 7.7 | 2/3 | 2.71s |
| #134 | GLM 5 Turbo none | Z.ai | 1 | 5.5 | 1/3 | 2.65s |
| #109 | GLM 5V Turbo none | Z.ai | 1 | 5.3 | 1/3 | 2.40s |
| #7 | Gemini 3.5 Flash medium | 1 | 7.7 | 2/3 | 2.38s | |
| #80 | Mimo V2 Omni medium | Xiaomi | 1 | 5.9 | 1/3 | 2.38s |