AI BENCHY Failures
Invalid tool call Failures
See which AI models run into Invalid tool call most often, so you can spot reliability risks before choosing one. Sort by: Failure Count ↑.
Categories
| Rank | Model | Company | Invalid tool call Count | Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #139 | DeepSeek V4 Flash none | DeepSeek | 1 | 5.0 | 5/21 | 26.8s |
| #145 | Laguna M.1 none | Poolside | 1 | 4.8 | 4/19 | 2.89s |
| #146 | Laguna Xs.2 none | Poolside | 1 | 4.8 | 5/19 | 806ms |
| #154 | Qwen3.5-9B none | Qwen | 1 | 4.6 | 4/21 | 1.89s |
| #158 | GLM 4.7 Flash medium | Z.ai | 1 | 4.4 | 4/21 | 35.1s |
| #159 | Ling-2.6-1T none | Inclusionai | 1 | 4.3 | 3/21 | 7.72s |
| #163 | Granite 4.1 8B none | IBM Granite | 1 | 4.0 | 2/21 | 728ms |
| #59 | GLM 5V Turbo medium | Z.ai | 2 | 7.2 | 11/21 | 23.1s |
| #138 | Ling-2.6-flash none | Inclusionai | 2 | 5.0 | 6/21 | 9.34s |