AI BENCHY Category Failures
Combined: Invalid tool call
Combined
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Combined, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
Categories
| Rank | Model | Company | Invalid tool call Count | Category Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #158 | GLM 4.7 Flash medium | Z.ai | 1 | 2.8 | 0/1 | 65.6s |
| #78 | Qwen3.6 27B medium | Qwen | 1 | 7.0 | 0/1 | 83.1s |
| #139 | DeepSeek V4 Flash none | DeepSeek | 1 | 4.5 | 0/1 | 112.0s |
| #133 | DeepSeek V3.2 none | DeepSeek | 1 | 6.5 | 0/1 | 115.9s |