Instructions following Ranking
See which AI models perform best on Instructions following, which ones stay reliable, and where the biggest gaps appear. Sort by: Total Cost ↓.
220/220
Filter models
No models match the current search and filters.
| Rank | Model | Company | Instructions following Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #154 | Owl Alpha none | Openrouter | 6.4 | 5.6 | $0.000 | 1/2 | 2.63s |
| #179 | North Mini Code none | Cohere | 6.5 | 5.1 | $0.000 | 1/2 | 30.7s |
| #184 | Qwen3.6 Plus Preview medium | Qwen | 6.5 | 4.9 | $0.000 | 1/2 | 3.40s |
| #193 | Hunter Alpha medium | OpenRouter | 9.9 | 4.7 | $0.000 | 2/2 | 4.18s |
| #197 | Cobuddy medium | Baidu | 9.8 | 4.7 | $0.000 | 2/2 | 11.6s |
| #203 | Elephant Alpha none | Openrouter | 9.8 | 4.3 | $0.000 | 2/2 | 1.03s |
| #205 | Elephant Alpha medium | Openrouter | 9.8 | 4.3 | $0.000 | 2/2 | 987ms |
| #206 | Hunter Alpha none | OpenRouter | 6.4 | 4.2 | $0.000 | 1/2 | 2.82s |
| #217 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 7.3 | 3.4 | $0.000 | 1/2 | 1.37s |
| #218 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 4.8 | 3.2 | $0.000 | 0/2 | 541ms |