Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster. Sort by: Tests Correct ↑.
Failure Reasons
Categories
80/80
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #180 | MiniMax M2.7 medium | Minimax | 1 | 3.8 | $0.163 | 0/2 | 72.1s |
| #181 | Laguna S 2.1 low | Poolside | 1 | 3.2 | $0.091 | 0/2 | 412.5s |
| #182 | DeepSeek V3.2 none | DeepSeek | 2 | 4.8 | $0.054 | 0/2 | 113.5s |
| #185 | GLM 4.7 Flash none | Z.ai | 2 | 3.0 | $0.016 | 0/2 | 50.2s |
| #187 | Ling-2.6-flash none | Inclusionai | 2 | 3.0 | $0.002 | 0/2 | 35.7s |
| #197 | Cobuddy medium | Baidu | 1 | 1.5 | $0.000 | 0/1 | 47.4s |
| #199 | MiniMax M2.5 medium | Minimax | 1 | 3.7 | $0.340 | 0/2 | 83.2s |
| #201 | Grok 4.20 Beta none | X AI | 1 | 1.5 | $0.087 | 0/1 | 6.48s |
| #202 | Laguna M.1 none | Poolside | 1 | 1.5 | $0.009 | 0/1 | 4.32s |
| #204 | GLM 4.7 Flash medium | Z.ai | 2 | 2.9 | $0.166 | 0/2 | 802.8s |
| #207 | Grok 4.20 none | X AI | 1 | 1.5 | $0.057 | 0/1 | 6.04s |
| #211 | Granite 4.1 8B none | IBM Granite | 2 | 3.0 | $0.007 | 0/2 | 9.28s |
| #4 | Gemini 3.5 Flash high | 1 | 8.2 | $1.976 | 1/2 | 84.1s | |
| #11 | Qwen3.7 Max medium | Qwen | 1 | 8.7 | $1.116 | 1/2 | 287.8s |
| #14 | Gemini 3.5 Flash low | 1 | 8.2 | $0.433 | 1/2 | 30.0s |