Coding: Did not follow instructions
Coding
Did not follow instructions
See which AI models are most likely to hit Did not follow instructions on Coding, so you can spot weak points faster. Sort by: Tests Correct ↓.
Failure Reasons
38/38
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #56 | Qwen3.8 27B high | Qwen | 1 | 0.8 | ~$0.298 | 2/3 | 248.6s |
| #64 | Claude Haiku 5.5 low | Anthropic | 1 | 0.8 | $0.061 | 2/3 | 7.30s |
| #107 | Gemini 3.5 Flash medium | 1 | 0.8 | $1.316 | 2/3 | 12.6s | |
| #160 | Qwen3.8 Flash Next xhigh | Qwen | 1 | 0.7 | ~$0.195 | 2/3 | 192.2s |
| #245 | Trinity Large Thinking medium | Arcee AI | 1 | 0.8 | $0.862 | 2/3 | 179.0s |
| #77 | DeepSeek V4 Flash 0731 high | DeepSeek | 1 | 0.6 | $0.576 | 1/3 | 252.7s |
| #108 | Qwen3.8 2.4T A95B high | Qwen | 1 | 0.6 | $2.949 | 1/3 | 160.5s |
| #131 | Seed-2.0-Code high | Bytedance Seed | 1 | 0.6 | $1.448 | 1/3 | 266.7s |
| #139 | DeepSeek V4 Pro 0423 none | DeepSeek | 1 | 0.6 | $0.083 | 1/3 | 13.4s |
| #178 | Claude Opus 4.8 none | Anthropic | 1 | 0.5 | $2.334 | 1/3 | 3.29s |
| #206 | Mercury 2.5 Preview high | Inception | 1 | 0.7 | $0.040 | 1/3 | 6.13s |
| #219 | Gemini 3.5 Flash minimal | 1 | 0.6 | $0.508 | 1/3 | 2.75s | |
| #231 | Ling 3.1 Flash low | Inclusionai | 1 | 0.6 | $0.000 | 1/3 | 107.3s |
| #232 | Claude Opus 4.6 medium | Anthropic | 1 | 0.6 | $2.961 | 1/3 | 30.1s |
| #266 | DeepSeek V4.1 Flash none | DeepSeek | 1 | 0.6 | $0.391 | 1/3 | 83.7s |