无效工具调用 失败
看看哪些 AI 模型最常遇到 无效工具调用,让你在选择前先发现稳定性风险。
131/131
筛选模型
没有模型匹配当前搜索和筛选条件。
| 排名 | 模型 | 公司 | 无效工具调用 次数 | 分数 | 总成本 | 测试正确 | 响应时间(平均) |
|---|---|---|---|---|---|---|---|
| #266 | Ling-2.6-flash none | Inclusionai | 3 | 4.9 | $0.002 | 6/22 | 10.7s |
| #59 | Inkling high | Thinkingmachines | 2 | 8.0 | $1.046 | 15/22 | 62.6s |
| #60 | Muse Spark 1.1 high | Meta | 2 | 8.0 | $2.253 | 12/22 | 36.9s |
| #153 | Gemini 3.5 Flash minimal | 2 | 6.7 | $0.300 | 13/22 | 2.65s | |
| #154 | GLM 5V Turbo medium | Z.ai | 2 | 6.7 | $0.457 | 11/21 | 23.1s |
| #160 | Inkling Small medium | Thinkingmachines | 2 | 6.6 | $0.113 | 10/22 | 6.30s |
| #165 | Qwen3.6 27B medium | Qwen | 2 | 6.5 | $1.045 | 10/22 | 106.1s |
| #174 | Ling-3.0-flash high | Inclusionai | 2 | 6.4 | $0.021 | 12/22 | 13.7s |
| #188 | Inkling low | Thinkingmachines | 2 | 6.1 | $0.187 | 10/22 | 5.11s |
| #189 | Trinity Large Thinking medium | Arcee AI | 2 | 6.1 | $0.756 | 8/22 | 85.4s |
| #190 | Qwen3.6 Flash none | Qwen | 2 | 6.1 | $0.062 | 7/22 | 3.73s |
| #194 | Trinity Large Thinking low | Arcee AI | 2 | 6.0 | $0.625 | 8/22 | 96.6s |
| #214 | Inkling Small low | Thinkingmachines | 2 | 5.7 | $0.055 | 9/22 | 2.07s |
| #229 | Qwen3.6 27B none | Qwen | 2 | 5.5 | $0.116 | 7/22 | 10.6s |
| #233 | DeepSeek V4 Flash 0423 none | DeepSeek | 2 | 5.4 | $0.037 | 4/22 | 36.2s |