AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category

Domain specific Ranking

See which AI models perform best on Domain specific, which ones stay reliable, and where the biggest gaps appear.

Models Shown

15

Average Domain specific Score

4.8

Rank Model Company Domain specific Score Score Tests Correct Response Time (avg)
#60 Gemma 4 26B A4B none Google 3.6 6.2 0/3 2.49s
#61 Seed-2.0-Lite none Bytedance Seed 3.6 6.2 0/3 1.33s
#64 DeepSeek V3.2 none DeepSeek 3.6 6.1 0/3 1.61s
#88 Nemotron 3 Super none NVIDIA 3.6 5.1 0/3 6.23s
#97 Qwen3.5-9B medium Qwen 3.6 4.4 0/3 137.7s
#13 GLM 5 medium Z.ai 3.5 8.4 0/3 0ms
#36 GPT-5.3 Chat none OpenAI 3.5 7.7 0/3 13.0s
#46 Kimi K2.5 medium Moonshot AI 3.5 7.0 0/3 137.3s
#86 GPT-5.4 Mini none OpenAI 3.5 5.1 0/3 937ms
#93 GLM 4.7 Flash medium Z.ai 3.5 4.6 0/3 174.6s
#9 Qwen3.6 Plus Preview medium Qwen 3.0 8.5 0/3 22.1s
#17 Gemini 3.1 Flash Lite Preview medium Google 3.0 8.2 0/3 4.21s
#35 MiMo-V2-Omni medium Xiaomi 3.0 7.7 0/3 55.1s
#37 Claude Opus 4.6 medium Anthropic 3.0 7.6 0/3 83.4s
#39 Seed-2.0-Mini medium Bytedance Seed 3.0 7.5 0/3 0ms

Top Models by Domain specific Score

Domain specific Score vs Total Cost

Top Models by Response Time (avg)