AI BENCHY
Advertise here

AI BENCHY Category

Tool Calling Ranking

See which AI models perform best on Tool Calling, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↓.

Models Shown

15

Average Tool Calling Score

8.7

Rank Model Company Tool Calling Score Score Tests Correct Response Time (avg)
#90 Gemini 3.1 Flash Lite none Google 10.0 6.4 1/1 2.97s
#91 GPT-5.5 none OpenAI 10.0 6.4 1/1 3.90s
#92 Laguna M.1 medium Poolside 10.0 6.4 1/1 6.31s
#93 Qwen3.6 Plus Preview medium Qwen 10.0 6.3 1/1 5.87s
#94 GPT-5 Nano medium OpenAI 10.0 6.3 1/1 33.3s
#95 Qwen3.5 Plus 2026-02-15 none Qwen 10.0 6.3 1/1 3.33s
#97 Gemini 2.5 Flash none Google 10.0 6.2 1/1 1.91s
#98 GLM 5 none Z.ai 10.0 6.1 1/1 11.1s
#99 gpt-oss-120b medium OpenAI 9.8 6.1 1/1 6.91s
#101 Mimo V2 Omni none Xiaomi 10.0 6.0 1/1 5.40s
#102 Gemma 4 26B A4B none Google 10.0 6.0 1/1 57.1s
#103 DeepSeek V4 Pro high DeepSeek 10.0 6.0 1/1 21.3s
#104 Nemotron 3 Ultra 550b A55b none NVIDIA 10.0 6.0 1/1 2.99s
#105 Nemotron 3 Super medium NVIDIA 10.0 5.8 1/1 39.7s
#106 Grok 4.20 Beta none X AI 10.0 5.8 1/1 4.79s

Top Models by Tool Calling Score

Tool Calling Score vs Total Cost

Top Models by Response Time (avg)