Coding: Did not follow instructions
Coding
Did not follow instructions
See which AI models are most likely to hit Did not follow instructions on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
38/38
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #139 | DeepSeek V4 Pro 0423 none | DeepSeek | 1 | 0.6 | $0.083 | 1/3 | 13.4s |
| #339 | DeepSeek V3.2 none | DeepSeek | 1 | 0.3 | $0.119 | 0/3 | 14.5s |
| #368 | Grok 4.1 Fast medium | X AI | 1 | 0.8 | $0.069 | 0/1 | 23.6s |
| #232 | Claude Opus 4.6 medium | Anthropic | 1 | 0.6 | $2.961 | 1/3 | 30.1s |
| #283 | Ling 3.1 Flash none | Inclusionai | 1 | 0.5 | $0.000 | 0/3 | 31.9s |
| #369 | Laguna M.1 medium | Poolside | 1 | 0.1 | $0.033 | 0/1 | 35.6s |
| #329 | Hy4 preview none | Tencent | 1 | 0.4 | $0.165 | 0/3 | 45.5s |
| #356 | Cobuddy medium | Baidu | 1 | 0.4 | $0.000 | 0/3 | 79.2s |
| #282 | Kimi K2.6 none | Moonshot AI | 1 | 0.6 | $0.486 | 1/3 | 82.6s |
| #266 | DeepSeek V4.1 Flash none | DeepSeek | 1 | 0.6 | $0.391 | 1/3 | 83.7s |
| #365 | Qwen3.5-9B medium | Qwen | 1 | 0.3 | $0.106 | 0/3 | 100.9s |
| #231 | Ling 3.1 Flash low | Inclusionai | 1 | 0.6 | $0.000 | 1/3 | 107.3s |
| #108 | Qwen3.8 2.4T A95B high | Qwen | 1 | 0.6 | $2.949 | 1/3 | 160.5s |
| #374 | Ling 3.0 Tiny low | Inclusionai | 1 | 0.3 | $0.000 | 0/3 | 170.7s |
| #373 | Ling 3.0 Tiny high | Inclusionai | 1 | 0.3 | $0.000 | 0/3 | 177.7s |