AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Failures

Wrong answer Failures

See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Tests Correct ↑.

Models Shown

15

Total Failures

572

Most Affected Model

LFM2-24B-A2B 9
Rank Model Company Wrong answer Count Score Tests Correct Response Time (avg)
#36 GPT-5.3 Chat none OpenAI 5 7.7 11/18 5.88s
#38 GPT-5.4 Nano medium OpenAI 4 7.6 11/18 11.2s
#39 Seed-2.0-Mini medium Bytedance Seed 2 7.5 11/18 69.7s
#40 GPT-5.2 medium OpenAI 2 7.5 11/18 14.0s
#41 MiMo-V2-Flash medium Xiaomi 3 7.5 11/18 23.4s
#42 Claude Sonnet 4.6 none Anthropic 3 7.4 11/18 4.98s
#30 Step 3.5 Flash medium Stepfun 3 7.9 11/17 26.8s
#18 GLM 5 Turbo medium Z.ai 3 8.1 12/18 17.7s
#23 MiMo-V2-Pro medium Xiaomi 3 8.1 12/18 12.3s
#25 Grok 4.20 Beta medium X AI 3 8.0 12/18 9.81s
#27 DeepSeek V3.2 medium DeepSeek 3 8.0 12/18 46.4s
#28 GPT-5.2 Chat none OpenAI 5 7.9 12/18 6.84s
#29 Gemini 3.1 Flash Lite Preview none Google 4 7.9 12/18 1.30s
#33 GLM 5.1 medium Z.ai 3 7.8 12/18 24.1s
#37 Claude Opus 4.6 medium Anthropic 2 7.6 12/18 21.1s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)