AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Failures

Wrong answer Failures

See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Score ↑.

Models Shown

15

Total Failures

572

Most Affected Model

LFM2-24B-A2B 9
Rank Model Company Wrong answer Count Score Tests Correct Response Time (avg)
#38 GPT-5.4 Nano medium OpenAI 4 7.6 11/18 11.2s
#37 Claude Opus 4.6 medium Anthropic 2 7.6 12/18 21.1s
#36 GPT-5.3 Chat none OpenAI 5 7.7 11/18 5.88s
#35 MiMo-V2-Omni medium Xiaomi 3 7.7 11/18 16.8s
#34 Kimi K2.6 medium Moonshot AI 2 7.7 11/18 45.2s
#33 GLM 5.1 medium Z.ai 3 7.8 12/18 24.1s
#32 Qwen3.5-Flash medium Qwen 1 7.8 11/18 66.7s
#31 GLM 5V Turbo medium Z.ai 3 7.8 11/18 15.0s
#30 Step 3.5 Flash medium Stepfun 3 7.9 11/17 26.8s
#29 Gemini 3.1 Flash Lite Preview none Google 4 7.9 12/18 1.30s
#28 GPT-5.2 Chat none OpenAI 5 7.9 12/18 6.84s
#27 DeepSeek V3.2 medium DeepSeek 3 8.0 12/18 46.4s
#25 Grok 4.20 Beta medium X AI 3 8.0 12/18 9.81s
#26 Claude Sonnet 4.6 medium Anthropic 2 8.0 13/18 12.7s
#24 Gemma 4 26B A4B medium Google 2 8.0 13/18 25.0s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)