AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category Failures

Anti-AI Tricks: Wrong answer

Anti-AI Tricks
Wrong answer

See which AI models are most likely to hit Wrong answer on Anti-AI Tricks, so you can spot weak points faster. Sort by: Response Time (avg) ↑.

Models Shown

15

Total Failures

245

Most Affected Model

Mistral Small 4 4
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#141 Nemotron 3 Super none NVIDIA 3 4.8 1/4 4.46s
#70 GPT-5.4 Nano medium OpenAI 1 8.3 3/4 4.52s
#79 Hunter Alpha medium OpenRouter 2 7.3 2/4 4.75s
#92 Laguna M.1 medium Poolside 1 6.5 2/4 4.87s
#122 GLM 4.7 Flash none Z.ai 3 5.2 1/4 5.51s
#135 Kimi K2.5 none Moonshot AI 4 3.6 0/4 6.24s
#100 Grok Build 0.1 none X AI 1 8.7 3/4 6.30s
#28 Gemini 2.5 Flash medium Google 1 8.4 3/4 6.30s
#159 Ling-2.6-1T none Inclusionai 4 3.4 0/4 6.55s
#42 GPT-5.2 medium OpenAI 1 6.5 2/4 7.81s
#150 Qwen3 Coder Next medium Qwen 3 3.5 0/4 8.64s
#133 DeepSeek V3.2 none DeepSeek 1 3.2 0/4 9.35s
#119 Cobuddy medium Baidu 1 8.7 3/4 10.00s
#99 gpt-oss-120b medium OpenAI 1 6.7 2/4 10.2s
#59 GLM 5V Turbo medium Z.ai 1 7.2 2/4 10.8s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost