AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category Failures

Puzzle Solving: Wrong answer

Puzzle Solving
Wrong answer

See which AI models are most likely to hit Wrong answer on Puzzle Solving, so you can spot weak points faster.

Models Shown

15

Total Failures

147

Most Affected Model

Qwen3.5-Flash 3
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#79 Hunter Alpha medium OpenRouter 1 6.1 1/3 5.35s
#80 Mimo V2 Omni medium Xiaomi 1 5.9 1/3 2.38s
#81 Mercury 2 medium Inception 1 5.4 1/3 949ms
#84 Grok 4.20 Multi Agent Beta medium X AI 1 6.7 1/3 5.19s
#85 Gemma 4 31B none Google 1 6.5 1/3 4.23s
#86 Grok 4.1 Fast medium X AI 1 5.3 1/3 7.40s
#88 Qwen3.7 Plus none Qwen 1 7.7 2/3 1.71s
#89 Hy3 preview low Tencent 1 5.3 1/3 7.51s
#90 Gemini 3.1 Flash Lite none Google 1 6.3 1/3 720ms
#91 GPT-5.5 none OpenAI 1 7.7 2/3 1.29s
#92 Laguna M.1 medium Poolside 1 5.3 1/3 10.2s
#94 GPT-5 Nano medium OpenAI 1 5.3 1/3 20.6s
#95 Qwen3.5 Plus 2026-02-15 none Qwen 1 7.7 2/3 2.71s
#97 Gemini 2.5 Flash none Google 1 7.7 2/3 604ms
#98 GLM 5 none Z.ai 1 7.7 2/3 1.91s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost