Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster. Sort by: Response Time (avg) ↓.
Failure Reasons
Categories
128/128
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #154 | Qwen3.6 Flash medium | Qwen | 1 | 6.5 | $0.741 | 1/2 | 299.2s |
| #209 | Trinity Large Thinking low | Arcee AI | 2 | 5.2 | $0.625 | 0/2 | 291.0s |
| #25 | Qwen3.7 Max medium | Qwen | 1 | 8.7 | $1.157 | 1/2 | 287.8s |
| #250 | Laguna S 2.1 medium | Poolside | 2 | 3.2 | $0.053 | 0/2 | 284.7s |
| #71 | Muse Glimmer 30B xhigh | Meta | 1 | 6.0 | $0.471 | 0/2 | 272.7s |
| #202 | Qwen3.5-Flash medium | Qwen | 1 | 6.4 | $0.142 | 1/2 | 266.6s |
| #118 | Qwen3.8 2.4T A95B high | Qwen | 1 | 6.9 | $2.357 | 1/2 | 266.1s |
| #289 | Trinity Large Thinking high | Arcee AI | 2 | 3.0 | $0.592 | 0/2 | 262.4s |
| #190 | Ring-2.6-1T medium | Inclusionai | 1 | 7.3 | $0.102 | 1/2 | 257.3s |
| #215 | Qwen3.5-Flash none | Qwen | 1 | 2.9 | $0.073 | 0/2 | 243.6s |
| #92 | Qwen3.7 Flash high | Qwen | 1 | 6.9 | $0.052 | 1/2 | 232.0s |
| #69 | Qwen3.8 2.4T A95B low | Qwen | 1 | 8.2 | $2.672 | 1/2 | 231.8s |
| #101 | Nemotron 3 Ultra medium | NVIDIA | 1 | 6.3 | $0.701 | 1/2 | 218.2s |
| #181 | Laguna XS 2.1 medium | Poolside | 1 | 6.3 | $0.068 | 1/2 | 218.1s |
| #255 | DeepSeek V4 Flash 0731 none | DeepSeek | 2 | 3.7 | $0.009 | 0/2 | 202.4s |