Did not follow instructions Failures
See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one. Sort by: Score ↑.
141/141
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #145 | GPT-5.4 none | OpenAI | 1 | 5.8 | $0.397 | 7/22 | 2.07s |
| #144 | Kimi K2.6 none | Moonshot AI | 3 | 5.8 | $0.184 | 7/22 | 19.6s |
| #142 | GPT-5.4 Mini none | OpenAI | 3 | 5.9 | $0.095 | 6/22 | 1.53s |
| #140 | Mimo V2 Omni medium | Xiaomi | 2 | 5.9 | $0.683 | 10/21 | 41.2s |
| #138 | GPT-5.6 Terra none | OpenAI | 1 | 6.0 | $0.349 | 8/22 | 1.65s |
| #137 | Grok 4.20 Beta medium | X AI | 1 | 6.0 | $0.750 | 14/18 | 9.75s |
| #136 | Step 3.5 Flash medium | Stepfun | 3 | 6.0 | $0.108 | 11/21 | 174.2s |
| #135 | Nemotron 3 Ultra none | NVIDIA | 1 | 6.1 | $0.095 | 8/22 | 3.87s |
| #134 | GPT-5 Nano medium | OpenAI | 2 | 6.1 | $0.114 | 9/22 | 54.9s |
| #133 | Qwen3.5-35B-A3B none | Qwen | 2 | 6.1 | $0.106 | 7/22 | 12.7s |
| #132 | Qwen3.5 Plus 2026-04-20 none | Qwen | 2 | 6.1 | $0.122 | 8/22 | 13.6s |
| #130 | Qwen3.6 Flash none | Qwen | 1 | 6.1 | $0.062 | 7/22 | 3.74s |
| #129 | Inkling low | Thinkingmachines | 2 | 6.1 | $0.187 | 10/22 | 5.15s |
| #128 | Gemini 3.1 Flash Lite none | 1 | 6.1 | $0.046 | 9/22 | 1.75s | |
| #127 | gpt-oss-120b medium | OpenAI | 3 | 6.1 | $0.019 | 9/22 | 21.9s |