AI BENCHY
Advertise here

AI BENCHY Failures

Did not follow instructions Failures

See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one. Sort by: Response Time (avg) ↓.

Models Shown

15

Total Failures

215

Most Affected Model

Kimi K2.5 2
Rank Model Company Did not follow instructions Count Score Tests Correct Response Time (avg)
#152 MiMo-V2-Flash none Xiaomi 2 4.6 4/21 2.76s
#101 Mimo V2 Omni none Xiaomi 1 6.0 8/21 2.44s
#120 Mimo V2 PRO none Xiaomi 2 5.6 7/21 2.27s
#104 Nemotron 3 Ultra 550b A55b none NVIDIA 1 6.0 8/21 2.27s
#81 Mercury 2 medium Inception 3 6.6 10/21 2.24s
#143 MiMo-V2.5 none Xiaomi 1 4.9 5/21 2.20s
#154 Qwen3.5-9B none Qwen 2 4.6 4/21 1.89s
#123 MiMo-V2.5-Pro none Xiaomi 4 5.5 6/21 1.78s
#147 GPT-4o-mini none OpenAI 1 4.8 5/21 1.77s
#115 Qwen3.5-27B none Qwen 2 5.7 7/21 1.68s
#157 Grok 4.1 Fast none X AI 3 4.4 3/19 1.62s
#128 Qwen3.6 Flash none Qwen 1 5.4 7/21 1.60s
#32 Gemini 3.5 Flash minimal Google 1 7.7 14/21 1.57s
#148 GPT-5.4 Nano none OpenAI 2 4.7 4/21 1.48s
#125 GPT-5.4 none OpenAI 1 5.5 7/21 1.42s

Top Models by Did not follow instructions Count

Did not follow instructions Count vs Score

Top Models by Response Time (avg)