Did not follow instructions Failures
See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one. Sort by: Response Time (avg) ↓.
215/215
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #58 | Gemini 2.5 Flash medium | 1 | 8.2 | $0.644 | 15/22 | 21.1s | |
| #189 | gpt-oss-120b medium | OpenAI | 3 | 6.1 | $0.020 | 9/22 | 20.8s |
| #244 | DeepSeek V4 Flash 0731 none | DeepSeek | 2 | 5.4 | $0.011 | 4/22 | 20.7s |
| #182 | MiMo-V2-Flash medium | Xiaomi | 1 | 6.3 | $0.043 | 12/21 | 20.1s |
| #42 | Muse Spark 1.2 medium | Meta | 3 | 8.6 | $1.339 | 14/22 | 19.7s |
| #222 | Kimi K2.6 none | Moonshot AI | 3 | 5.7 | $0.233 | 6/22 | 19.6s |
| #261 | Qwen3.5-9B none | Qwen | 2 | 5.1 | $0.021 | 4/22 | 19.2s |
| #128 | Muse Glimmer 30B low | Meta | 3 | 7.1 | $0.143 | 11/22 | 18.6s |
| #267 | DeepSeek V3.2 none | DeepSeek | 1 | 5.0 | $0.054 | 6/22 | 17.9s |
| #313 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 1 | 3.4 | $0.000 | 4/19 | 17.1s |
| #30 | GPT-5.3-Codex medium | OpenAI | 2 | 8.9 | $0.988 | 16/22 | 16.9s |
| #164 | Gemini 3.1 Flash Lite Preview low | 1 | 6.6 | $0.646 | 14/22 | 16.7s | |
| #172 | Hy3 preview medium | Tencent | 1 | 6.5 | $0.049 | 14/21 | 16.3s |
| #59 | Muse Spark 1.3 low | Meta | 1 | 8.2 | $0.689 | 15/22 | 16.0s |
| #284 | Laguna M.1 medium | Poolside | 1 | 4.7 | $0.033 | 9/19 | 14.7s |