Invalid tool call Failure Ranking

See which AI models run into Invalid tool call most often, so you can spot reliability risks before choosing one. Sort by: Tests Correct ↓.

Models Shown

Total Failures

100

Most Affected Model

Gemini 3.5 Flash 1

Categories

In category Combined91 In category Tool Calling9

83/83

Rank	Model	Company	Invalid tool call Count	Score	Total Cost	Tests Correct	Response Time (avg)
#110	Gemma 4 31B medium	Google	1	6.3	$0.163	14/22	75.4s
Total Tests 22 Wrong Tests 8 Total Cost $0.163 Response Time (avg) 75.4s
#24	Muse Spark 1.1 low	Meta	1	8.3	$0.647	13/22	11.5s
Total Tests 22 Wrong Tests 9 Total Cost $0.647 Response Time (avg) 11.5s
#45	DeepSeek V4 Flash high	DeepSeek	1	7.7	$0.042	13/22	49.7s
Total Tests 22 Wrong Tests 9 Total Cost $0.042 Response Time (avg) 49.7s
#51	Nemotron 3 Ultra medium	NVIDIA	1	7.5	$0.774	13/22	32.2s
Total Tests 22 Wrong Tests 9 Total Cost $0.774 Response Time (avg) 32.2s
#55	GPT-5.6 Terra low	OpenAI	1	7.5	$0.519	13/22	5.31s
Total Tests 22 Wrong Tests 9 Total Cost $0.519 Response Time (avg) 5.31s
#58	Qwen3.5-27B medium	Qwen	1	7.4	$1.627	13/22	111.9s
Total Tests 22 Wrong Tests 9 Total Cost $1.627 Response Time (avg) 111.9s
#64	Gemini 3.1 Flash Lite Preview medium	Google	1	7.3	$0.115	13/22	4.61s
Total Tests 22 Wrong Tests 9 Total Cost $0.115 Response Time (avg) 4.61s
#65	Gemini 3.1 Flash Lite medium	Google	1	7.3	$0.117	13/22	4.27s
Total Tests 22 Wrong Tests 9 Total Cost $0.117 Response Time (avg) 4.27s
#90	Qwen3.6 35B A3B medium	Qwen	1	6.7	$0.746	13/22	58.1s
Total Tests 22 Wrong Tests 9 Total Cost $0.746 Response Time (avg) 58.1s
#104	Gemini 3.1 Flash Lite Preview low	Google	1	6.5	$0.646	13/22	16.7s
Total Tests 22 Wrong Tests 9 Total Cost $0.646 Response Time (avg) 16.7s
#27	Muse Spark 1.1 high	Meta	2	8.1	$1.694	12/22	31.5s
Total Tests 22 Wrong Tests 10 Total Cost $1.694 Response Time (avg) 31.5s
#56	GPT-5.4 Mini medium	OpenAI	1	7.5	$0.756	12/22	25.9s
Total Tests 22 Wrong Tests 10 Total Cost $0.756 Response Time (avg) 25.9s
#67	Step 3.7 Flash low	Stepfun	1	7.3	$0.454	12/22	20.7s
Total Tests 22 Wrong Tests 10 Total Cost $0.454 Response Time (avg) 20.7s
#68	Kimi K2.6 medium	Moonshot AI	1	7.2	$1.036	12/22	110.0s
Total Tests 22 Wrong Tests 10 Total Cost $1.036 Response Time (avg) 110.0s
#75	Grok 4.20 medium	X AI	1	7.1	$0.777	12/22	29.5s
Total Tests 22 Wrong Tests 10 Total Cost $0.777 Response Time (avg) 29.5s

←

1 2 3 4 5 6

→

Invalid tool call Failures

Filter models

Top Models by Invalid tool call Count

Invalid tool call Count vs Score

Top Models by Response Time (avg)