AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category

Tool Calling Ranking

See which AI models perform best on Tool Calling, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.

Models Shown

15

Average Tool Calling Score

8.7

Rank Model Company Tool Calling Score Score Tests Correct Response Time (avg)
#66 GPT-5.4 none OpenAI 10.0 5.9 1/1 2.75s
#67 Qwen3.5-27B none Qwen 10.0 5.9 1/1 3.54s
#69 Kimi K2.6 none Moonshot AI 10.0 5.8 1/1 4.46s
#70 Qwen3.5-122B-A10B none Qwen 10.0 5.7 1/1 2.04s
#71 MiniMax M2.5 medium Minimax 10.0 5.7 1/1 15.4s
#72 Hunter Alpha none OpenRouter 10.0 5.7 1/1 6.02s
#73 Mistral Small 4 medium Mistral 10.0 5.7 1/1 3.50s
#75 GLM 5.1 none Z.ai 10.0 5.6 1/1 10.7s
#76 Kimi K2.5 none Moonshot AI 10.0 5.5 1/1 14.0s
#77 GLM 5 Turbo none Z.ai 10.0 5.5 1/1 8.21s
#78 Trinity Large Preview none Arcee AI 10.0 5.3 1/1 6.67s
#79 Grok 4.20 Beta none X AI 10.0 5.3 1/1 4.79s
#82 Grok 4.20 none X AI 10.0 5.2 1/1 4.63s
#83 Mistral Small 4 none Mistral 10.0 5.2 1/1 1.40s
#87 Qwen3 Coder Next none Qwen 10.0 5.1 1/1 2.47s

Top Models by Tool Calling Score

Tool Calling Score vs Total Cost

Top Models by Response Time (avg)