AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category Failures

Puzzle Solving: API error

Puzzle Solving
API error

See which AI models are most likely to hit API error on Puzzle Solving, so you can spot weak points faster.

Models Shown

12

Total Failures

13

Most Affected Model

Qwen3.6 Plus Preview 2
Rank Model Company API error Count Category Score Tests Correct Response Time (avg)
#93 Qwen3.6 Plus Preview medium Qwen 2 5.3 1/3 7.52s
#82 Hy3 preview high Tencent 1 7.7 2/3 27.9s
#89 Hy3 preview low Tencent 1 5.3 1/3 7.51s
#92 Laguna M.1 medium Poolside 1 5.3 1/3 10.2s
#103 DeepSeek V4 Pro high DeepSeek 1 5.9 1/3 34.8s
#107 Laguna Xs.2 medium Poolside 1 5.3 1/3 1.93s
#133 DeepSeek V3.2 none DeepSeek 1 7.6 2/3 6.91s
#145 Laguna M.1 none Poolside 1 3.0 0/3 891ms
#146 Laguna Xs.2 none Poolside 1 5.3 1/3 650ms
#149 Nemotron 3 Nano Omni 30b A3b Reasoning medium NVIDIA 1 2.9 0/3 1.40s
#160 LFM2-24B-A2B none Liquid 1 3.8 0/3 1.78s
#162 Nemotron 3 Nano Omni 30b A3b Reasoning none NVIDIA 1 3.0 0/3 532ms

Top Models by API error Count

API error Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost