AI BENCHY
Your ad here

AI BENCHY Category

Domain specific Ranking

See which AI models perform best on Domain specific, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↑.

Models Shown

15

Average Domain specific Score

4.8

Rank Model Company Domain specific Score Score Tests Correct Response Time (avg)
#9 Qwen3.6 Plus Preview medium Qwen 3.0 8.5 0/3 22.1s
#13 GLM 5 medium Z.ai 3.5 8.4 0/3 0ms
#17 Gemini 3.1 Flash Lite Preview medium Google 3.0 8.2 0/3 4.21s
#18 GLM 5 Turbo medium Z.ai 2.9 8.1 0/3 71.1s
#19 Qwen3.5-122B-A10B medium Qwen 2.9 8.1 0/3 63.4s
#20 Qwen3.6 Plus medium Qwen 2.9 8.1 0/3 29.6s
#24 Gemma 4 26B A4B medium Google 2.9 8.0 0/3 23.6s
#26 Claude Sonnet 4.6 medium Anthropic 2.9 8.0 0/3 0ms
#35 MiMo-V2-Omni medium Xiaomi 3.0 7.7 0/3 55.1s
#36 GPT-5.3 Chat none OpenAI 3.5 7.7 0/3 13.0s
#37 Claude Opus 4.6 medium Anthropic 3.0 7.6 0/3 83.4s
#39 Seed-2.0-Mini medium Bytedance Seed 3.0 7.5 0/3 0ms
#43 Qwen3.5-35B-A3B medium Qwen 4.1 7.4 0/3 88.3s
#44 GPT-5.4 Mini medium OpenAI 4.1 7.3 0/3 65.3s
#45 GPT-5 Mini medium OpenAI 3.6 7.0 0/3 44.6s

Top Models by Domain specific Score

Domain specific Score vs Total Cost

Top Models by Response Time (avg)