Agentic: Invalid tool call
Agentic
Invalid tool call
See which AI models are most likely to hit Invalid tool call on Agentic, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
Categories
44/44
Filter models
No models match the current search and filters.
| Rank | Model | Company | Invalid tool call Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #226 | Space Bunny Alpha xhigh | Stealth | 1 | 5.2 | $0.000 | 0/1 | 159.1s |
| #283 | Laguna XS 2.1 none | Poolside | 1 | 5.0 | $0.019 | 0/1 | 70.8s |
| #293 | Mercury 2.5 low | Inception | 1 | 5.0 | $0.020 | 0/1 | 47.1s |
| #314 | Mercury 2.5 Preview none | Inception | 1 | 3.9 | $0.023 | 0/1 | 85.0s |
| #267 | GPT-6 Luna none | OpenAI | 1 | 5.0 | $0.026 | 0/1 | 51.1s |
| #261 | Solar Pro 4 none | Upstage | 1 | 5.0 | $0.033 | 0/1 | 69.2s |
| #171 | Mercury 2.5 Preview medium | Inception | 1 | 3.9 | $0.037 | 0/1 | 74.6s |
| #336 | Granite 4.2 8B none | IBM Granite | 1 | 5.0 | $0.047 | 0/1 | 49.2s |
| #106 | GPT-6 Luna medium | OpenAI | 1 | 7.4 | $0.048 | 0/1 | 71.7s |
| #308 | MiMo-V2.5 none | Xiaomi | 1 | 3.9 | $0.063 | 0/1 | 124.2s |
| #110 | GPT-6 Luna high | OpenAI | 1 | 5.2 | $0.074 | 0/1 | 82.8s |
| #200 | Laguna XS 2.1 medium | Poolside | 1 | 6.1 | $0.083 | 0/1 | 66.4s |
| #122 | GPT-5.6 Luna medium | OpenAI | 1 | 7.4 | $0.101 | 0/1 | 58.1s |
| #212 | LongCat 2.0 none | Meituan | 1 | 6.1 | $0.104 | 0/1 | 81.8s |
| #284 | Qwen3.6 35B A3B none | Qwen | 1 | 5.2 | $0.105 | 0/1 | 65.6s |