AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Kategori AI BENCHY

Peringkat Pemanggilan alat

Lihat model AI mana yang paling baik di Pemanggilan alat, mana yang tetap andal, dan di mana kesenjangan terbesar muncul. Urutkan berdasarkan: Metrik ↑.

Model yang ditampilkan

15

Rata-rata Skor Pemanggilan alat

8.7

Model terbaik

Grok 4.1 Fast 2.8
Peringkat Model Perusahaan Skor Pemanggilan alat Skor Tes benar Waktu respons (rata-rata)
#44 GPT-5.4 Mini medium OpenAI 4.7 7.3 0/1 9.62s
#80 MiniMax M2.7 medium Minimax 4.7 5.3 0/1 12.0s
#88 Nemotron 3 Super none NVIDIA 4.7 5.1 0/1 16.0s
#31 GLM 5V Turbo medium Z.ai 7.0 7.8 0/1 12.5s
#68 gpt-oss-120b medium OpenAI 9.8 5.8 1/1 6.91s
#1 Gemini 3 Flash Preview medium Google 10.0 10.0 1/1 10.6s
#2 Gemini 3.1 Pro Preview medium Google 10.0 9.6 1/1 23.1s
#3 Claude Opus 4.7 medium Anthropic 10.0 9.2 1/1 4.17s
#4 Claude Opus 4.7 none Anthropic 10.0 9.2 1/1 4.74s
#5 Gemini 3 Flash Preview low Google 10.0 8.8 1/1 4.99s
#6 Seed-2.0-Lite medium Bytedance Seed 10.0 8.6 1/1 12.4s
#7 GPT-5.3-Codex medium OpenAI 10.0 8.6 1/1 6.37s
#8 Qwen3.5 Plus 2026-02-15 medium Qwen 10.0 8.5 1/1 7.54s
#9 Qwen3.6 Plus Preview medium Qwen 10.0 8.5 1/1 5.87s
#10 Qwen3.5-27B medium Qwen 10.0 8.4 1/1 7.45s

Model teratas menurut Skor Pemanggilan alat

Skor Pemanggilan alat vs total biaya

Model teratas menurut Waktu respons (rata-rata)