API error Failure Ranking

See which AI models run into API error most often, so you can spot reliability risks before choosing one. Sort by: Response Time (avg) ↓.

Models Shown

Total Failures

161

Most Affected Model

Step 3.5 Flash 1

Categories

In category Coding45 In category Combined26 In category Tool Calling17 In category Anti-AI Tricks14 In category Data parsing and extraction14 In category Trivia13 In category General Intelligence12 In category Puzzle Solving12 In category Domain specific7 In category Instructions following1

68/68

Rank	Model	Company	API error Count	Score	Total Cost	Tests Correct	Response Time (avg)
#130	Step 3.5 Flash medium	Stepfun	1	6.0	$0.108	11/21	174.2s
Total Tests 21 Wrong Tests 10 Total Cost $0.108 Response Time (avg) 174.2s
#137	North Mini Code medium	Cohere	1	5.9	$0.000	9/22	137.1s
Total Tests 22 Wrong Tests 13 Total Cost $0.000 Response Time (avg) 137.1s
#60	LongCat 2.0 medium	Meituan	1	7.4	$0.478	12/22	136.6s
Total Tests 22 Wrong Tests 10 Total Cost $0.478 Response Time (avg) 136.6s
#33	Kimi K3 max	Moonshot AI	2	8.0	$3.112	16/22	122.5s
Total Tests 22 Wrong Tests 6 Total Cost $3.112 Response Time (avg) 122.5s
#119	Qwen3.5-35B-A3B medium	Qwen	1	6.2	$0.837	11/22	112.5s
Total Tests 22 Wrong Tests 11 Total Cost $0.837 Response Time (avg) 112.5s
#91	LongCat 2.0 low	Meituan	1	6.7	$0.391	10/22	100.3s
Total Tests 22 Wrong Tests 12 Total Cost $0.391 Response Time (avg) 100.3s
#57	Qwen3.5 Plus 2026-02-15 medium	Qwen	1	7.5	$0.437	14/22	89.2s
Total Tests 22 Wrong Tests 8 Total Cost $0.437 Response Time (avg) 89.2s
#114	Qwen3.5-Flash medium	Qwen	1	6.2	$0.139	12/22	84.8s
Total Tests 22 Wrong Tests 10 Total Cost $0.139 Response Time (avg) 84.8s
#52	Kimi K2.7 Code medium	Moonshot AI	1	7.5	$0.751	12/22	84.2s
Total Tests 22 Wrong Tests 10 Total Cost $0.751 Response Time (avg) 84.2s
#204	Qwen3.5-9B medium	Qwen	1	3.8	$0.036	3/22	82.2s
Total Tests 22 Wrong Tests 19 Total Cost $0.036 Response Time (avg) 82.2s
#46	DeepSeek V4 Pro high	DeepSeek	1	7.7	$0.200	10/22	79.1s
Total Tests 22 Wrong Tests 12 Total Cost $0.200 Response Time (avg) 79.1s
#110	Gemma 4 31B medium	Google	2	6.3	$0.163	14/22	75.4s
Total Tests 22 Wrong Tests 8 Total Cost $0.163 Response Time (avg) 75.4s
#108	Ring-2.6-1T medium	Inclusionai	2	6.3	$0.103	11/22	68.7s
Total Tests 22 Wrong Tests 11 Total Cost $0.103 Response Time (avg) 68.7s
#76	DeepSeek V3.2 medium	DeepSeek	2	7.0	$0.078	11/22	68.6s
Total Tests 22 Wrong Tests 11 Total Cost $0.078 Response Time (avg) 68.6s
#90	Qwen3.6 35B A3B medium	Qwen	2	6.7	$0.746	13/22	58.1s
Total Tests 22 Wrong Tests 9 Total Cost $0.746 Response Time (avg) 58.1s

1 2 3 4 5

→

API error Failures

Filter models

Top Models by API error Count

API error Count vs Score

Top Models by Response Time (avg)