Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster.
Failure Reasons
Categories
125/125
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #60 | Muse Spark 1.1 high | Meta | 2 | 5.9 | $2.253 | 0/2 | 70.3s |
| #153 | Gemini 3.5 Flash minimal | 2 | 3.0 | $0.300 | 0/2 | 14.4s | |
| #160 | Inkling Small medium | Thinkingmachines | 2 | 3.8 | $0.113 | 0/2 | 8.89s |
| #165 | Qwen3.6 27B medium | Qwen | 2 | 6.7 | $1.045 | 0/2 | 584.1s |
| #174 | Ling-3.0-flash high | Inclusionai | 2 | 2.9 | $0.021 | 0/2 | 50.4s |
| #188 | Inkling low | Thinkingmachines | 2 | 2.9 | $0.187 | 0/2 | 22.7s |
| #189 | Trinity Large Thinking medium | Arcee AI | 2 | 2.9 | $0.756 | 0/2 | 382.7s |
| #190 | Qwen3.6 Flash none | Qwen | 2 | 3.8 | $0.062 | 0/2 | 26.5s |
| #194 | Trinity Large Thinking low | Arcee AI | 2 | 5.2 | $0.625 | 0/2 | 291.0s |
| #229 | Qwen3.6 27B none | Qwen | 2 | 3.2 | $0.116 | 0/2 | 83.1s |
| #233 | DeepSeek V4 Flash 0423 none | DeepSeek | 2 | 4.6 | $0.037 | 0/2 | 179.6s |
| #234 | Laguna S 2.1 medium | Poolside | 2 | 3.2 | $0.053 | 0/2 | 284.7s |
| #239 | DeepSeek V4 Flash 0731 none | DeepSeek | 2 | 3.7 | $0.011 | 0/2 | 202.4s |
| #240 | Laguna S 2.1 high | Poolside | 2 | 2.9 | $0.114 | 0/2 | 702.3s |
| #256 | Qwen3.5-9B none | Qwen | 2 | 3.0 | $0.021 | 0/2 | 194.0s |