Tool Calling Ranking
See which AI models perform best on Tool Calling, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↓.
330/330
Filter models
No models match the current search and filters.
| Rank | Model | Company | Tool Calling Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #155 | Step 3.7 Flash high | Stepfun | 10.0 | 6.8 | $1.229 | 1/1 | 2.79s |
| #168 | Gemini 3.5 Flash minimal | 10.0 | 6.7 | $0.300 | 1/1 | 2.79s | |
| #305 | Elephant Alpha none | Openrouter | 3.0 | 4.3 | $0.000 | 0/1 | 2.79s |
| #246 | Qwen3.8 27B none | Qwen | 10.0 | 5.5 | ~$0.006 | 1/1 | 2.75s |
| #224 | GPT-5.4 none | OpenAI | 10.0 | 5.8 | $0.397 | 1/1 | 2.75s |
| #271 | Mercury 2.5 low | Inception | 9.8 | 5.1 | $0.011 | 1/1 | 2.72s |
| #256 | Laguna S 2.1 high | Poolside | 10.0 | 5.4 | $0.114 | 1/1 | 2.69s |
| #300 | Qwen3 Coder Next medium | Qwen | 10.0 | 4.6 | $0.034 | 1/1 | 2.64s |
| #203 | Inkling low | Thinkingmachines | 3.0 | 6.1 | $0.187 | 0/1 | 2.57s |
| #280 | GPT-4o-mini none | OpenAI | 10.0 | 5.0 | $0.010 | 1/1 | 2.51s |
| #264 | Inkling none | Thinkingmachines | 3.0 | 5.2 | $0.147 | 0/1 | 2.50s |
| #205 | Qwen3.6 Flash none | Qwen | 10.0 | 6.1 | $0.062 | 1/1 | 2.49s |
| #270 | Qwen3 Coder Next none | Qwen | 10.0 | 5.1 | $0.026 | 1/1 | 2.47s |
| #272 | MiMo-V2.5 none | Xiaomi | 10.0 | 5.1 | $0.025 | 1/1 | 2.43s |
| #109 | Mercury 2.5 Preview medium | Inception | 3.0 | 7.5 | $0.024 | 0/1 | 2.37s |