AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category Failures

Puzzle Solving: Wrong answer

Puzzle Solving
Wrong answer

See which AI models are most likely to hit Wrong answer on Puzzle Solving, so you can spot weak points faster. Sort by: Response Time (avg) ↓.

Models Shown

15

Total Failures

147

Most Affected Model

Qwen3.6 27B 1
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#24 GPT-5.2 Chat none OpenAI 1 7.7 2/3 4.10s
#135 Kimi K2.5 none Moonshot AI 3 3.0 0/3 4.04s
#64 MiMo-V2-Flash medium Xiaomi 1 7.7 2/3 3.87s
#70 GPT-5.4 Nano medium OpenAI 2 4.1 0/3 3.79s
#116 Hunter Alpha none OpenRouter 1 5.8 1/3 3.71s
#41 Nemotron 3 Ultra 550b A55b medium NVIDIA 2 5.5 1/3 3.54s
#111 Owl Alpha medium Openrouter 1 5.3 1/3 3.40s
#28 Gemini 2.5 Flash medium Google 1 7.7 2/3 3.18s
#105 Nemotron 3 Super medium NVIDIA 2 3.0 0/3 3.15s
#110 Seed-2.0-Lite none Bytedance Seed 2 5.3 1/3 2.78s
#95 Qwen3.5 Plus 2026-02-15 none Qwen 1 7.7 2/3 2.71s
#134 GLM 5 Turbo none Z.ai 1 5.5 1/3 2.65s
#109 GLM 5V Turbo none Z.ai 1 5.3 1/3 2.40s
#7 Gemini 3.5 Flash medium Google 1 7.7 2/3 2.38s
#80 Mimo V2 Omni medium Xiaomi 1 5.9 1/3 2.38s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost