AI BENCHY
Your ad here

AI BENCHY Category

Domain specific Ranking

See which AI models perform best on Domain specific, which ones stay reliable, and where the biggest gaps appear.

Models Shown

15

Average Domain specific Score

4.8

Rank Model Company Domain specific Score Score Tests Correct Response Time (avg)
#50 Hunter Alpha medium OpenRouter 3.0 6.7 0/3 10.5s
#53 GLM 5 none Z.ai 3.0 6.6 0/3 2.24s
#67 Qwen3.5-27B none Qwen 3.0 5.9 0/3 540ms
#79 Grok 4.20 Beta none X AI 3.0 5.3 0/3 611ms
#80 MiniMax M2.7 medium Minimax 3.0 5.3 0/3 19.0s
#81 Elephant medium Openrouter 3.0 5.2 0/3 925ms
#82 Grok 4.20 none X AI 3.0 5.2 0/3 687ms
#84 gpt-oss-120b none OpenAI 3.0 5.2 0/3 35.0s
#85 Elephant none Openrouter 3.0 5.2 0/3 927ms
#89 GPT-4o-mini none OpenAI 3.0 4.9 0/3 637ms
#90 Qwen3.5-9B none Qwen 3.0 4.8 0/3 464ms
#19 Qwen3.5-122B-A10B medium Qwen 2.9 8.1 0/3 63.4s
#20 Qwen3.6 Plus medium Qwen 2.9 8.1 0/3 29.6s
#26 Claude Sonnet 4.6 medium Anthropic 2.9 8.0 0/3 0ms
#54 Mercury 2 medium Inception 2.9 6.5 0/3 6.48s

Top Models by Domain specific Score

Domain specific Score vs Total Cost

Top Models by Response Time (avg)