Did not follow instructions Failure Ranking

See which AI models run into Did not follow instructions most often, so you can spot reliability risks before choosing one. Sort by: Response Time (avg) ↑.

Models Shown

Total Failures

245

Most Affected Model

Nemotron 3 Nano Omni 30b A3b Reasoning 2

Categories

In category Puzzle Solving90 In category General Intelligence78 In category Anti-AI Tricks33 In category Instructions following18 In category Coding16 In category Tool Calling8 In category Combined1 In category Domain specific1

140/140

Rank	Model	Company	Did not follow instructions Count	Score	Total Cost	Tests Correct	Response Time (avg)
#129	Nemotron 3 Ultra none	NVIDIA	1	6.1	$0.095	8/22	3.87s
Total Tests 22 Wrong Tests 14 Total Cost $0.095 Response Time (avg) 3.87s
#154	MiMo-V2.5-Pro none	Xiaomi	4	5.5	$0.068	6/22	4.12s
Total Tests 22 Wrong Tests 16 Total Cost $0.068 Response Time (avg) 4.12s
#65	Gemini 3.1 Flash Lite medium	Google	1	7.3	$0.117	13/22	4.27s
Total Tests 22 Wrong Tests 9 Total Cost $0.117 Response Time (avg) 4.27s
#64	Gemini 3.1 Flash Lite Preview medium	Google	1	7.3	$0.115	13/22	4.61s
Total Tests 22 Wrong Tests 9 Total Cost $0.115 Response Time (avg) 4.61s
#168	MiMo-V2.5 none	Xiaomi	1	5.1	$0.025	5/22	4.62s
Total Tests 22 Wrong Tests 17 Total Cost $0.025 Response Time (avg) 4.62s
#196	Hunter Alpha none	OpenRouter	2	4.2	$0.000	6/18	4.70s
Total Tests 18 Wrong Tests 12 Total Cost $0.000 Response Time (avg) 4.70s
#103	Qwen3.5-27B none	Qwen	2	6.5	$0.090	8/22	4.76s
Total Tests 22 Wrong Tests 14 Total Cost $0.090 Response Time (avg) 4.76s
#66	Claude Opus 4.8 none	Anthropic	1	7.3	$1.166	13/22	4.91s
Total Tests 22 Wrong Tests 9 Total Cost $1.166 Response Time (avg) 4.91s
#117	GPT-5.6 Luna low	OpenAI	1	6.2	$0.249	10/22	5.04s
Total Tests 22 Wrong Tests 12 Total Cost $0.249 Response Time (avg) 5.04s
#123	Inkling low	Thinkingmachines	2	6.1	$0.187	10/22	5.15s
Total Tests 22 Wrong Tests 12 Total Cost $0.187 Response Time (avg) 5.15s
#115	Gemma 4 31B none	Google	1	6.2	$0.035	10/22	5.34s
Total Tests 22 Wrong Tests 12 Total Cost $0.035 Response Time (avg) 5.34s
#161	Qwen3.6 35B A3B none	Qwen	2	5.3	$0.061	4/22	5.52s
Total Tests 22 Wrong Tests 18 Total Cost $0.061 Response Time (avg) 5.52s
#177	Nemotron 3 Super none	NVIDIA	2	4.9	$0.008	5/22	5.97s
Total Tests 22 Wrong Tests 17 Total Cost $0.008 Response Time (avg) 5.97s
#112	Claude Sonnet 5 none	Anthropic	1	6.3	$0.548	8/22	6.04s
Total Tests 22 Wrong Tests 14 Total Cost $0.548 Response Time (avg) 6.04s
#54	GPT-5.3 Chat none	OpenAI	2	7.5	$0.571	13/22	6.88s
Total Tests 22 Wrong Tests 9 Total Cost $0.571 Response Time (avg) 6.88s

Did not follow instructions Failures

Filter models

Top Models by Did not follow instructions Count

Did not follow instructions Count vs Score

Top Models by Response Time (avg)