AI BENCHY
Advertise here

AI BENCHY Category

Coding Ranking

See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear.

Models Shown

15

Average Coding Score

7.5

Rank Model Company Coding Score Score Tests Correct Response Time (avg)
#20 GLM 5 Turbo medium Z.ai 10.0 8.1 1/1 12.3s
#21 Qwen3.6 35B A3B medium Qwen 10.0 8.0 1/1 32.6s
#22 Hy3 preview high Tencent 10.0 8.0 1/1 99.8s
#23 Gemini 3.1 Flash Lite Preview medium Google 10.0 8.0 1/1 4.34s
#24 Grok 4.3 medium X AI 10.0 8.0 1/1 45.7s
#25 Gemini 2.5 Flash medium Google 10.0 7.9 1/1 16.2s
#26 GPT-5.4 medium OpenAI 10.0 7.9 1/1 13.0s
#27 Gemini 3.1 Flash Lite medium Google 10.0 7.9 1/1 3.26s
#29 Gemini 3 Flash Preview none Google 10.0 7.9 1/1 1.59s
#30 Gemini 3.1 Flash Lite Preview low Google 10.0 7.9 1/1 2.20s
#32 MiMo-V2.5 medium Xiaomi 10.0 7.8 1/1 31.5s
#34 Hy3 preview medium Tencent 10.0 7.8 1/1 31.4s
#35 Claude Sonnet 4.6 medium Anthropic 10.0 7.8 1/1 35.8s
#37 MiMo-V2-Pro medium Xiaomi 10.0 7.7 1/1 52.1s
#39 Hy3 preview low Tencent 10.0 7.7 1/1 27.9s

Top Models by Coding Score

Coding Score vs Total Cost

Top Models by Response Time (avg)