Did not follow instructions Failures
See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one.
215/215
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #212 | GPT-5.4 Mini none | OpenAI | 3 | 5.9 | $0.095 | 6/22 | 1.58s |
| #217 | Nemotron 3 Super medium | NVIDIA | 3 | 5.7 | $0.050 | 8/22 | 52.3s |
| #222 | Kimi K2.6 none | Moonshot AI | 3 | 5.7 | $0.233 | 6/22 | 19.6s |
| #224 | Gemini 3.1 Flash Lite high | 3 | 5.6 | $2.044 | 10/18 | 62.0s | |
| #230 | Owl Alpha none | Openrouter | 3 | 5.6 | $0.000 | 7/21 | 9.88s |
| #255 | Nemotron 3.5 Lightning medium | NVIDIA | 3 | 5.1 | $0.156 | 5/22 | 70.1s |
| #281 | Trinity Large Preview none | Arcee AI | 3 | 4.8 | $0.008 | 4/21 | 2.98s |
| #285 | Cobuddy medium | Baidu | 3 | 4.7 | $0.000 | 7/21 | 39.9s |
| #288 | Qwen3 Coder Next medium | Qwen | 3 | 4.6 | $0.034 | 3/22 | 9.07s |
| #289 | MiniMax M2.5 medium | Minimax | 3 | 4.6 | $0.327 | 5/22 | 69.1s |
| #293 | Elephant Alpha none | Openrouter | 3 | 4.3 | $0.000 | 5/21 | 1.22s |
| #306 | Ling 3.0 Tiny high | Inclusionai | 3 | 3.8 | $0.000 | 4/22 | 75.7s |
| #308 | Grok 4.1 Fast none | X AI | 3 | 3.8 | $0.008 | 3/19 | 1.62s |
| #30 | GPT-5.3-Codex medium | OpenAI | 2 | 8.9 | $0.988 | 16/22 | 16.9s |
| #31 | Muse Spark 1.3 high | Meta | 2 | 8.9 | $1.653 | 16/22 | 52.5s |