Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster.
Failure Reasons
Categories
77/77
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #121 | gpt-oss-120b medium | OpenAI | 1 | 6.5 | $0.019 | 1/2 | 24.0s |
| #125 | Qwen3.5-Flash none | Qwen | 1 | 2.9 | $0.073 | 0/2 | 243.6s |
| #127 | Qwen3.5-35B-A3B none | Qwen | 1 | 3.8 | $0.106 | 0/2 | 128.3s |
| #132 | GPT-5.6 Terra none | OpenAI | 1 | 2.9 | $0.349 | 0/2 | 7.02s |
| #137 | North Mini Code medium | Cohere | 1 | 2.9 | $0.000 | 0/2 | 554.9s |
| #142 | Qwen3.5-122B-A10B none | Qwen | 1 | 5.2 | $0.247 | 0/2 | 129.3s |
| #151 | GLM 5.1 none | Z.ai | 1 | 2.8 | $0.164 | 0/2 | 46.9s |
| #156 | Gemma 4 26B A4B none | 1 | 3.0 | $0.015 | 0/2 | 37.2s | |
| #159 | GPT-5.6 Luna none | OpenAI | 1 | 3.2 | $0.142 | 0/2 | 6.68s |
| #160 | Laguna XS 2.1 none | Poolside | 1 | 3.0 | $0.008 | 0/2 | 10.4s |
| #164 | Inkling none | Thinkingmachines | 1 | 2.9 | $0.147 | 0/2 | 25.7s |
| #172 | MiniMax M2.7 medium | Minimax | 1 | 3.8 | $0.163 | 0/2 | 72.1s |
| #188 | Cobuddy medium | Baidu | 1 | 1.5 | $0.000 | 0/1 | 47.4s |
| #190 | MiniMax M2.5 medium | Minimax | 1 | 3.7 | $0.340 | 0/2 | 83.2s |
| #191 | Grok 4.20 Beta none | X AI | 1 | 1.5 | $0.087 | 0/1 | 6.48s |