Instructions following: Did not follow instructions
Instructions following
Did not follow instructions
See which AI models are most likely to hit Did not follow instructions on Instructions following, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
18/18
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #207 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 1 | 7.3 | $0.000 | 1/2 | 1.37s |
| #208 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 1 | 4.8 | $0.000 | 0/2 | 541ms |
| #201 | Granite 4.1 8B none | IBM Granite | 1 | 3.6 | $0.007 | 0/2 | 344ms |
| #203 | Grok 4.1 Fast none | X AI | 1 | 3.0 | $0.008 | 0/2 | 685ms |
| #183 | Trinity Large Preview none | Arcee AI | 1 | 3.5 | $0.008 | 0/2 | 822ms |
| #140 | Nemotron 3 Super medium | NVIDIA | 1 | 7.3 | $0.050 | 1/2 | 6.97s |
| #185 | Grok 4.1 Fast medium | X AI | 1 | 6.5 | $0.069 | 1/2 | 4.63s |
| #130 | Step 3.5 Flash medium | Stepfun | 1 | 8.3 | $0.108 | 1/2 | 4.78s |
| #172 | MiniMax M2.7 medium | Minimax | 1 | 3.8 | $0.163 | 0/2 | 12.8s |
| #46 | DeepSeek V4 Pro high | DeepSeek | 1 | 7.8 | $0.200 | 1/2 | 8.73s |
| #117 | GPT-5.6 Luna low | OpenAI | 1 | 8.5 | $0.249 | 1/2 | 2.04s |
| #190 | MiniMax M2.5 medium | Minimax | 1 | 7.5 | $0.340 | 1/2 | 621ms |
| #132 | GPT-5.6 Terra none | OpenAI | 1 | 8.5 | $0.349 | 1/2 | 1.15s |
| #83 | GPT-5.6 Sol none | OpenAI | 1 | 8.5 | $0.524 | 1/2 | 1.33s |
| #24 | Muse Spark 1.1 low | Meta | 1 | 7.3 | $0.647 | 1/2 | 5.42s |