AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Failures

Did not follow instructions Failures

See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one. Sort by: Score ↑.

Models Shown

15

Total Failures

215

Most Affected Model

Granite 4.1 8B 4
Rank Model Company Did not follow instructions Count Score Tests Correct Response Time (avg)
#163 Granite 4.1 8B none IBM Granite 4 4.0 2/21 728ms
#162 Nemotron 3 Nano Omni 30b A3b Reasoning none NVIDIA 2 4.1 2/19 728ms
#161 Qwen3.5-9B medium Qwen 1 4.2 3/21 82.2s
#160 LFM2-24B-A2B none Liquid 1 4.2 2/16 782ms
#159 Ling-2.6-1T none Inclusionai 2 4.3 3/21 7.72s
#158 GLM 4.7 Flash medium Z.ai 2 4.4 4/21 35.1s
#157 Grok 4.1 Fast none X AI 3 4.4 3/19 1.62s
#156 Hy3 preview none Tencent 4 4.4 4/21 12.9s
#155 Mercury 2 none Inception 1 4.5 4/21 653ms
#154 Qwen3.5-9B none Qwen 2 4.6 4/21 1.89s
#153 Qwen3.6 35B A3B none Qwen 2 4.6 4/21 3.73s
#152 MiMo-V2-Flash none Xiaomi 2 4.6 4/21 2.76s
#151 Trinity Large Preview none Arcee AI 3 4.6 4/21 2.98s
#150 Qwen3 Coder Next medium Qwen 3 4.6 4/21 8.58s
#149 Nemotron 3 Nano Omni 30b A3b Reasoning medium NVIDIA 1 4.6 4/19 17.1s

Top Models by Did not follow instructions Count

Did not follow instructions Count vs Score

Top Models by Response Time (avg)