AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category Failures

Coding: Wrong answer

Coding
Wrong answer

See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster.

Models Shown

15

Total Failures

120

Most Affected Model

Qwen3.6 Flash 2
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#134 Nemotron 3 Super none NVIDIA 2 3.4 0/2 3.02s
#135 Mistral Small 4 none Mistral 2 4.0 0/2 1.03s
#138 GPT-4o-mini none OpenAI 2 3.2 0/2 2.05s
#142 Qwen3.5-9B none Qwen 2 4.4 0/2 5.39s
#143 Mercury 2 none Inception 2 3.5 0/2 831ms
#147 GPT-5.4 Nano none OpenAI 2 5.4 0/2 1.09s
#1 Gemini 3 Flash Preview medium Google 1 7.9 1/2 96.0s
#3 Gemini 3.5 Flash low Google 1 6.8 1/2 5.54s
#4 Gemini 3.1 Pro Preview medium Google 1 7.0 1/2 54.3s
#9 Gemini 3.5 Flash none Google 1 8.2 1/2 39.6s
#11 GPT-5.5 medium OpenAI 1 8.2 1/2 69.7s
#12 Gemini 3 Flash Preview low Google 1 7.3 1/2 6.66s
#14 Qwen3.6 Max Preview medium Qwen 1 8.2 1/2 178.0s
#20 Qwen3.5 Plus 2026-02-15 medium Qwen 1 7.6 1/2 193.8s
#21 Seed-2.0-Lite medium Bytedance Seed 1 7.0 1/2 107.7s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost