AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category Failures

Anti-AI Tricks: Wrong answer

Anti-AI Tricks
Wrong answer

See which AI models are most likely to hit Wrong answer on Anti-AI Tricks, so you can spot weak points faster. Sort by: Response Time (avg) ↑.

Models Shown

15

Total Failures

245

Most Affected Model

Mistral Small 4 4
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#156 Hy3 preview none Tencent 1 4.8 1/4 11.1s
#60 Kimi K2.6 medium Moonshot AI 1 7.0 2/4 11.6s
#138 Ling-2.6-flash none Inclusionai 1 6.8 2/4 11.8s
#78 Qwen3.6 27B medium Qwen 1 8.3 3/4 12.6s
#54 GPT-5 Mini medium OpenAI 1 7.1 2/4 13.9s
#113 DeepSeek V4 Pro none DeepSeek 3 3.5 0/4 14.0s
#67 MiniMax M3 medium Minimax 2 5.5 1/4 14.9s
#158 GLM 4.7 Flash medium Z.ai 2 4.7 1/4 15.0s
#103 DeepSeek V4 Pro high DeepSeek 1 6.4 2/4 16.5s
#19 Seed-2.0-Lite medium Bytedance Seed 1 8.3 3/4 18.0s
#139 DeepSeek V4 Flash none DeepSeek 4 3.0 0/4 20.2s
#94 GPT-5 Nano medium OpenAI 2 6.5 2/4 25.5s
#31 DeepSeek V4 Flash high DeepSeek 1 8.3 3/4 28.5s
#126 gpt-oss-120b none OpenAI 1 6.5 2/4 32.8s
#161 Qwen3.5-9B medium Qwen 1 5.1 1/4 34.4s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost