AI BENCHY
Advertise here

AI BENCHY Category Failures

Puzzle Solving: Did not follow instructions

Puzzle Solving
Did not follow instructions

See which AI models are most likely to hit Did not follow instructions on Puzzle Solving, so you can spot weak points faster. Sort by: Response Time (avg) ↑.

Models Shown

15

Total Failures

78

Most Affected Model

Mistral Small 4 1
Rank Model Company Did not follow instructions Count Category Score Tests Correct Response Time (avg)
#109 GLM 5V Turbo none Z.ai 1 5.3 1/3 2.40s
#134 GLM 5 Turbo none Z.ai 1 5.5 1/3 2.65s
#105 Nemotron 3 Super medium NVIDIA 1 3.0 0/3 3.15s
#111 Owl Alpha medium Openrouter 1 5.3 1/3 3.40s
#116 Hunter Alpha none OpenRouter 1 5.8 1/3 3.71s
#70 GPT-5.4 Nano medium OpenAI 1 4.1 0/3 3.79s
#121 Owl Alpha none Openrouter 1 5.4 1/3 4.18s
#85 Gemma 4 31B none Google 1 6.5 1/3 4.23s
#45 GPT-5.4 Mini medium OpenAI 1 7.8 2/3 4.37s
#156 Hy3 preview none Tencent 1 3.1 0/3 4.56s
#15 GPT-5.3-Codex medium OpenAI 1 9.0 2/3 5.05s
#51 Mimo V2 PRO medium Xiaomi 1 6.4 1/3 5.08s
#118 Qwen3.6 27B none Qwen 1 5.3 1/3 5.15s
#84 Grok 4.20 Multi Agent Beta medium X AI 1 6.7 1/3 5.19s
#23 GLM 5 Turbo medium Z.ai 1 8.7 2/3 5.23s

Top Models by Did not follow instructions Count

Did not follow instructions Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost