AI BENCHY
Advertise here

AI BENCHY Category

Combined Ranking

See which AI models perform best on Combined, which ones stay reliable, and where the biggest gaps appear.

Models Shown

15

Average Combined Score

6.3

Rank Model Company Combined Score Score Tests Correct Response Time (avg)
#67 MiniMax M3 medium Minimax 10.0 7.1 1/1 65.3s
#69 Claude Opus 4.6 medium Anthropic 10.0 7.0 1/1 76.7s
#71 Step 3.7 Flash high Stepfun 10.0 7.0 1/1 13.0s
#72 DeepSeek V3.2 medium DeepSeek 10.0 7.0 1/1 93.1s
#73 Seed-2.0-Mini medium Bytedance Seed 10.0 6.9 1/1 262.8s
#75 Ring-2.6-1T medium Inclusionai 10.0 6.9 1/1 304.2s
#76 Kimi K2.5 medium Moonshot AI 10.0 6.8 1/1 71.4s
#80 Mimo V2 Omni medium Xiaomi 10.0 6.7 1/1 25.9s
#81 Mercury 2 medium Inception 10.0 6.6 1/1 3.28s
#82 Hy3 preview high Tencent 10.0 6.6 1/1 113.1s
#86 Grok 4.1 Fast medium X AI 10.0 6.5 1/1 37.6s
#88 Qwen3.7 Plus none Qwen 10.0 6.4 1/1 29.4s
#89 Hy3 preview low Tencent 10.0 6.4 1/1 78.7s
#93 Qwen3.6 Plus Preview medium Qwen 10.0 6.3 1/1 35.0s
#94 GPT-5 Nano medium OpenAI 10.0 6.3 1/1 66.0s

Top Models by Combined Score

Combined Score vs Total Cost

Top Models by Response Time (avg)