AI BENCHY Failures
Did not follow instructions Failures
See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one. Sort by: Total Cost ↑.
131/131
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #132 | Owl Alpha medium | Openrouter | 2 | 5.8 | $0.000 | 8/21 | 11.9s |
| #134 | Owl Alpha none | Openrouter | 3 | 5.8 | $0.000 | 7/21 | 9.88s |
| #158 | North Mini Code none | Cohere | 2 | 5.1 | $0.000 | 4/21 | 29.8s |
| #159 | Hunter Alpha medium | OpenRouter | 2 | 5.1 | $0.000 | 8/18 | 10.3s |
| #167 | Cobuddy medium | Baidu | 3 | 4.9 | $0.000 | 7/21 | 39.9s |
| #180 | Elephant Alpha none | Openrouter | 3 | 4.6 | $0.000 | 5/21 | 1.22s |
| #181 | Elephant Alpha medium | Openrouter | 2 | 4.5 | $0.000 | 6/21 | 1.27s |
| #182 | Hunter Alpha none | OpenRouter | 2 | 4.5 | $0.000 | 6/18 | 4.70s |
| #194 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 1 | 3.6 | $0.000 | 4/19 | 17.1s |
| #195 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 2 | 3.5 | $0.000 | 2/19 | 728ms |
| #197 | LFM2-24B-A2B none | Liquid | 1 | 2.4 | $0.001 | 2/16 | 782ms |
| #171 | Ling-2.6-flash none | Inclusionai | 2 | 4.9 | $0.001 | 6/21 | 9.34s |
| #120 | Gemma 4 31B none | 1 | 6.1 | $0.002 | 10/21 | 4.05s | |
| #186 | Hy3 preview none | Tencent | 4 | 4.3 | $0.003 | 4/21 | 12.9s |
| #191 | Granite 4.1 8B none | IBM Granite | 4 | 4.0 | $0.003 | 2/21 | 728ms |