Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster.
Failure Reasons
Categories
128/128
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #43 | GLM 5.3 Flash max | Z.ai | 1 | 7.3 | $0.071 | 1/2 | 138.7s |
| #45 | Muse Spark 1.2 high | Meta | 1 | 8.0 | $1.771 | 1/2 | 39.9s |
| #46 | Muse Spark 1.2 medium | Meta | 1 | 8.5 | $1.339 | 1/2 | 24.9s |
| #47 | Muse Spark 1.1 medium | Meta | 1 | 8.3 | $1.420 | 1/2 | 42.6s |
| #48 | Claude Fable 5 medium | Anthropic | 1 | 6.5 | $3.004 | 1/2 | 27.5s |
| #64 | Muse Spark 1.1 low | Meta | 1 | 6.6 | $0.673 | 1/2 | 29.4s |
| #67 | Muse Spark 1.3 low | Meta | 1 | 7.3 | $0.689 | 1/2 | 25.7s |
| #68 | Claude Sonnet 5 medium | Anthropic | 1 | 7.3 | $0.922 | 1/2 | 51.9s |
| #69 | Qwen3.8 2.4T A95B low | Qwen | 1 | 8.2 | $2.672 | 1/2 | 231.8s |
| #71 | Muse Glimmer 30B xhigh | Meta | 1 | 6.0 | $0.471 | 0/2 | 272.7s |
| #72 | Inkling high | Thinkingmachines | 1 | 7.3 | $1.046 | 1/2 | 63.8s |
| #78 | GPT-5.6 Terra high | OpenAI | 1 | 8.7 | $0.836 | 1/2 | 13.7s |
| #81 | Step 3.7 Flash medium | Stepfun | 1 | 7.3 | $0.495 | 1/2 | 80.9s |
| #83 | Qwen3.7 Plus medium | Qwen | 1 | 8.2 | $0.277 | 1/2 | 190.3s |
| #85 | Gemini 3.5 Flash Lite medium | 1 | 7.3 | $0.272 | 1/2 | 17.9s |