AI BENCHY Category
Tool Calling Ranking
See which AI models perform best on Tool Calling, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↑.
| Rank | Model | Company | Tool Calling Score | Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #32 | Gemini 3.5 Flash minimal | 10.0 | 7.7 | 1/1 | 2.79s | |
| #33 | Hy3 preview medium | Tencent | 10.0 | 7.7 | 1/1 | 15.0s |
| #34 | Qwen3.7 Max none | Qwen | 10.0 | 7.7 | 1/1 | 3.92s |
| #35 | Gemini 3 PRO Preview medium | 10.0 | 7.6 | 1/1 | 12.0s | |
| #36 | Qwen3.5 Plus 2026-04-20 medium | Qwen | 10.0 | 7.6 | 1/1 | 14.7s |
| #37 | Gemma 4 26B A4B medium | 10.0 | 7.6 | 1/1 | 9.01s | |
| #38 | Grok 4.3 medium | X AI | 10.0 | 7.6 | 1/1 | 17.7s |
| #39 | Qwen3.6 Flash medium | Qwen | 10.0 | 7.5 | 1/1 | 4.00s |
| #40 | Gemini 3.1 Flash Lite Preview medium | 10.0 | 7.5 | 1/1 | 3.80s | |
| #41 | Nemotron 3 Ultra 550b A55b medium | NVIDIA | 10.0 | 7.5 | 1/1 | 7.72s |
| #43 | MiMo-V2.5-Pro medium | Xiaomi | 10.0 | 7.5 | 1/1 | 16.9s |
| #44 | Gemini 3.1 Flash Lite medium | 10.0 | 7.5 | 1/1 | 4.55s | |
| #47 | Grok Build 0.1 medium | X AI | 10.0 | 7.4 | 1/1 | 13.1s |
| #48 | Gemini 3 Flash Preview none | 10.0 | 7.4 | 1/1 | 3.35s | |
| #49 | Qwen3.5-Flash medium | Qwen | 10.0 | 7.4 | 1/1 | 10.3s |