AI BENCHY
Advertise here

AI BENCHY Category Failures

Instructions following: Wrong answer

Instructions following
Wrong answer

See which AI models are most likely to hit Wrong answer on Instructions following, so you can spot weak points faster. Sort by: Response Time (avg) ↑.

Models Shown

8

Total Failures

53

Most Affected Model

Granite 4.1 8B 1
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#159 Ling-2.6-1T none Inclusionai 1 6.4 1/2 5.36s
#55 GLM 5.1 medium Z.ai 1 6.4 1/2 7.47s
#150 Qwen3 Coder Next medium Qwen 1 6.3 1/2 7.49s
#140 Qwen3 Coder Next none Qwen 1 6.3 1/2 7.78s
#113 DeepSeek V4 Pro none DeepSeek 1 6.3 1/2 8.23s
#108 Qwen3.5-Flash none Qwen 1 6.3 1/2 8.81s
#111 Owl Alpha medium Openrouter 1 6.5 1/2 10.2s
#130 MiniMax M2.7 medium Minimax 1 3.8 0/2 12.8s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost