Agentic: Invalid tool call
Agentic
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Agentic, so you can spot weak points faster.
Failure Reasons
Categories
44/44
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #157 | Ember-1 low | Fireworks | 1 | 4.7 | $1.488 | 0/1 | 121.3s |
| #158 | Qwen3.5-122B-A10B medium | Qwen | 1 | 6.1 | $1.152 | 0/1 | 95.3s |
| #161 | Muse Glimmer 30B low | Meta | 1 | 6.1 | $0.266 | 0/1 | 105.0s |
| #164 | Ember-1 high | Fireworks | 1 | 6.1 | $2.405 | 0/1 | 88.3s |
| #171 | Mercury 2.5 Preview medium | Inception | 1 | 3.9 | $0.037 | 0/1 | 74.6s |
| #175 | Grok 4.7 xhigh | X AI | 1 | 3.0 | $4.054 | 0/1 | 304.8s |
| #176 | Grok Build 0.1 medium | X AI | 1 | 3.4 | $1.419 | 0/1 | 69.0s |
| #179 | Grok 4.7 high | X AI | 1 | 3.0 | $3.582 | 0/1 | 322.2s |
| #195 | GLM 5.3 low | Z.ai | 1 | 5.2 | $0.473 | 0/1 | 51.3s |
| #199 | Grok 4.3 medium | X AI | 1 | 3.9 | $0.989 | 0/1 | 51.6s |
| #200 | Laguna XS 2.1 medium | Poolside | 1 | 6.1 | $0.083 | 0/1 | 66.4s |
| #212 | LongCat 2.0 none | Meituan | 1 | 6.1 | $0.104 | 0/1 | 81.8s |
| #217 | Qwen3.5 Plus 2026-04-20 none | Qwen | 1 | 7.4 | $0.226 | 0/1 | 133.3s |
| #225 | Gemini 3.5 Flash none | 1 | 3.5 | $1.778 | 0/1 | 89.3s | |
| #226 | Space Bunny Alpha xhigh | Stealth | 1 | 5.2 | $0.000 | 0/1 | 159.1s |