AI BENCHY
Advertise here

AI BENCHY Category Failures

Anti-AI Tricks: Wrong answer

Anti-AI Tricks
Wrong answer

See which AI models are most likely to hit Wrong answer on Anti-AI Tricks, so you can spot weak points faster. Sort by: Response Time (avg) ↓.

Models Shown

15

Total Failures

245

Most Affected Model

Seed-2.0-Mini 1
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#54 GPT-5 Mini medium OpenAI 1 7.1 2/4 13.9s
#78 Qwen3.6 27B medium Qwen 1 8.3 3/4 12.6s
#138 Ling-2.6-flash none Inclusionai 1 6.8 2/4 11.8s
#60 Kimi K2.6 medium Moonshot AI 1 7.0 2/4 11.6s
#156 Hy3 preview none Tencent 1 4.8 1/4 11.1s
#59 GLM 5V Turbo medium Z.ai 1 7.2 2/4 10.8s
#99 gpt-oss-120b medium OpenAI 1 6.7 2/4 10.2s
#119 Cobuddy medium Baidu 1 8.7 3/4 10.00s
#133 DeepSeek V3.2 none DeepSeek 1 3.2 0/4 9.35s
#150 Qwen3 Coder Next medium Qwen 3 3.5 0/4 8.64s
#42 GPT-5.2 medium OpenAI 1 6.5 2/4 7.81s
#159 Ling-2.6-1T none Inclusionai 4 3.4 0/4 6.55s
#28 Gemini 2.5 Flash medium Google 1 8.4 3/4 6.30s
#100 Grok Build 0.1 none X AI 1 8.7 3/4 6.30s
#135 Kimi K2.5 none Moonshot AI 4 3.6 0/4 6.24s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost