Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster.
Failure Reasons
Categories
125/125
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #241 | Inkling Small none | Thinkingmachines | 1 | 6.5 | $0.054 | 1/2 | 9.33s |
| #242 | Laguna XS 2.1 none | Poolside | 1 | 3.0 | $0.008 | 0/2 | 10.4s |
| #247 | Dots 3 Note Preview none | Dots Studio | 1 | 3.0 | $0.000 | 0/2 | 14.8s |
| #248 | Inkling none | Thinkingmachines | 1 | 2.9 | $0.147 | 0/2 | 25.7s |
| #250 | Nemotron 3.5 Lightning medium | NVIDIA | 1 | 6.2 | $0.156 | 1/2 | 136.9s |
| #259 | MiniMax M2.7 medium | Minimax | 1 | 3.8 | $0.208 | 0/2 | 72.1s |
| #261 | Laguna S 2.1 low | Poolside | 1 | 3.2 | $0.082 | 0/2 | 412.5s |
| #265 | Mercury 2.5 Preview none | Inception | 1 | 3.0 | $0.018 | 0/2 | 31.3s |
| #270 | Nemotron 3.5 Lightning low | NVIDIA | 1 | 6.3 | $0.140 | 1/2 | 142.4s |
| #280 | Cobuddy medium | Baidu | 1 | 1.5 | $0.000 | 0/1 | 47.4s |
| #282 | Granite 4.2 8B high | IBM Granite | 1 | 5.0 | $0.084 | 0/2 | 570.4s |
| #284 | MiniMax M2.5 medium | Minimax | 1 | 3.7 | $0.327 | 0/2 | 83.2s |
| #286 | Grok 4.20 Beta none | X AI | 1 | 1.5 | $0.087 | 0/1 | 6.48s |
| #287 | Laguna M.1 none | Poolside | 1 | 1.5 | $0.009 | 0/1 | 4.32s |
| #293 | Grok 4.20 none | X AI | 1 | 1.5 | $0.057 | 0/1 | 6.04s |