Agentic: Did not follow instructions
Agentic
Did not follow instructions
See which AI models are most likely to hit Did not follow instructions on Agentic, so you can spot weak points faster. Sort by: Failure Count ↑.
Failure Reasons
14/14
Filter models
No models match the current search and filters.
| Rank | Model | Company | Did not follow instructions Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #66 | Gemini 3.5 Flash Lite high | 1 | 7.4 | $0.649 | 0/1 | 63.8s | |
| #104 | Step 3.7 Flash medium | Stepfun | 1 | 7.4 | $0.571 | 0/1 | 93.9s |
| #132 | GLM 5.3 Flash high | Z.ai | 1 | 7.4 | $0.084 | 0/1 | 111.5s |
| #135 | Gemini 2.5 Flash medium | 1 | 3.0 | $0.644 | 0/1 | 6.54s | |
| #147 | Gemini 3.8 Flash low | 1 | 5.9 | $0.334 | 0/1 | 62.9s | |
| #186 | Qwen3.7 Flash medium | Qwen | 1 | 5.2 | $0.076 | 0/1 | 149.0s |
| #197 | GLM 5.2 none | Z.ai | 1 | 5.2 | $0.161 | 0/1 | 52.3s |
| #203 | Qwen3.5 Plus 2026-02-15 none | Qwen | 1 | 5.9 | $0.173 | 0/1 | 183.5s |
| #213 | Qwen3.5-35B-A3B medium | Qwen | 1 | 7.4 | $1.249 | 0/1 | 126.8s |
| #232 | Trinity Large Thinking low | Arcee AI | 1 | 6.4 | $0.703 | 0/1 | 73.0s |
| #266 | Gemini 2.5 Flash none | 1 | 3.0 | $0.017 | 0/1 | 1.81s | |
| #317 | Laguna S 2.1 none | Poolside | 1 | 5.4 | $0.047 | 0/1 | 149.3s |
| #324 | Laguna S 2.1 low | Poolside | 1 | 2.8 | $0.109 | 0/1 | 152.1s |
| #338 | Mercury 2 none | Inception | 1 | 3.0 | $0.033 | 0/1 | 9.00s |