Coding: Did not follow instructions
Coding
Did not follow instructions
See which AI models are most likely to hit Did not follow instructions on Coding, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
38/38
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #365 | Qwen3.5-9B medium | Qwen | 1 | 0.3 | $0.106 | 0/3 | 100.9s |
| #339 | DeepSeek V3.2 none | DeepSeek | 1 | 0.3 | $0.119 | 0/3 | 14.5s |
| #313 | MiMo-V2.5-Pro none | Xiaomi | 1 | 0.4 | $0.135 | 0/3 | 1.41s |
| #329 | Hy4 preview none | Tencent | 1 | 0.4 | $0.165 | 0/3 | 45.5s |
| #160 | Qwen3.8 Flash Next xhigh | Qwen | 1 | 0.7 | ~$0.195 | 2/3 | 192.2s |
| #248 | Qwen3.5 Plus 2026-04-20 none | Qwen | 1 | 0.4 | $0.216 | 0/3 | 1.69s |
| #198 | GLM 5.2 none | Z.ai | 1 | 0.4 | $0.250 | 0/3 | 7.55s |
| #56 | Qwen3.8 27B high | Qwen | 1 | 0.8 | ~$0.298 | 2/3 | 248.6s |
| #266 | DeepSeek V4.1 Flash none | DeepSeek | 1 | 0.6 | $0.391 | 1/3 | 83.7s |
| #282 | Kimi K2.6 none | Moonshot AI | 1 | 0.6 | $0.486 | 1/3 | 82.6s |
| #219 | Gemini 3.5 Flash minimal | 1 | 0.6 | $0.508 | 1/3 | 2.75s | |
| #77 | DeepSeek V4 Flash 0731 high | DeepSeek | 1 | 0.6 | $0.576 | 1/3 | 252.7s |
| #269 | Trinity Large Thinking high | Arcee AI | 1 | 0.4 | $0.794 | 0/3 | 245.0s |
| #245 | Trinity Large Thinking medium | Arcee AI | 1 | 0.8 | $0.862 | 2/3 | 179.0s |
| #177 | Claude Sonnet 5.5 low | Anthropic | 2 | 0.3 | $0.909 | 0/3 | 5.81s |