Invalid tool call Failures
See which AI models run into Invalid tool call most often, so you can spot reliability risks before choosing one. Sort by: Tests Correct ↑.
Categories
83/83
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #173 | DeepSeek V3.2 none | DeepSeek | 2 | 5.0 | $0.054 | 6/22 | 18.3s |
| #176 | GLM 4.7 Flash none | Z.ai | 2 | 4.9 | $0.016 | 6/22 | 9.15s |
| #178 | Ling-2.6-flash none | Inclusionai | 3 | 4.9 | $0.002 | 6/22 | 10.7s |
| #195 | Elephant Alpha medium | Openrouter | 1 | 4.3 | $0.000 | 6/21 | 1.27s |
| #198 | Laguna Xs.2 medium | Poolside | 1 | 4.1 | $0.015 | 6/19 | 6.73s |
| #124 | Qwen3.6 Flash none | Qwen | 2 | 6.1 | $0.062 | 7/22 | 3.74s |
| #127 | Qwen3.5-35B-A3B none | Qwen | 1 | 6.1 | $0.106 | 7/22 | 12.7s |
| #151 | GLM 5.1 none | Z.ai | 1 | 5.5 | $0.164 | 7/22 | 6.70s |
| #152 | Qwen3.6 27B none | Qwen | 2 | 5.5 | $0.087 | 7/22 | 10.7s |
| #188 | Cobuddy medium | Baidu | 1 | 4.7 | $0.000 | 7/21 | 39.9s |
| #191 | Grok 4.20 Beta none | X AI | 1 | 4.4 | $0.087 | 6/18 | 1.19s |
| #197 | Grok 4.20 none | X AI | 1 | 4.1 | $0.057 | 6/18 | 1.11s |
| #125 | Qwen3.5-Flash none | Qwen | 1 | 6.1 | $0.073 | 8/22 | 25.3s |
| #132 | GPT-5.6 Terra none | OpenAI | 1 | 6.0 | $0.349 | 8/22 | 1.65s |
| #156 | Gemma 4 26B A4B none | 1 | 5.5 | $0.015 | 8/22 | 7.64s |