AI BENCHY
Your ad here

AI BENCHY Category Failures

Instructions following: Wrong answer

Instructions following
Wrong answer

See which AI models are most likely to hit Wrong answer on Instructions following, so you can spot weak points faster. Sort by: Response Time (avg) ↑.

Models Shown

9

Total Failures

44

Most Affected Model

Mistral Small 4 1
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#93 GLM 4.7 Flash medium Z.ai 1 6.2 1/2 2.97s
#36 GPT-5.3 Chat none OpenAI 1 8.3 1/2 3.29s
#55 MiMo-V2-Omni none Xiaomi 1 6.5 1/2 4.18s
#28 GPT-5.2 Chat none OpenAI 1 7.5 1/2 5.46s
#92 Qwen3 Coder Next medium Qwen 1 4.8 0/2 7.34s
#33 GLM 5.1 medium Z.ai 1 6.4 1/2 7.47s
#87 Qwen3 Coder Next none Qwen 2 4.8 0/2 7.71s
#59 Qwen3.5-Flash none Qwen 1 6.3 1/2 8.81s
#80 MiniMax M2.7 medium Minimax 1 3.7 0/2 12.6s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost