Instructions following: Wrong answer
Instructions following
Wrong answer
See which AI models are most likely to hit Wrong answer on Instructions following, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
61/61
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #63 | Claude Sonnet 4.6 none | Anthropic | 1 | 6.5 | $0.661 | 1/2 | 1.96s |