AI BENCHY
Advertise here

AI BENCHY Category

Coding Ranking

See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↑.

Models Shown

15

Average Coding Score

6.1

Rank Model Company Coding Score Score Tests Correct Response Time (avg)
#73 GPT-5 Mini medium OpenAI 10.0 6.9 2/2 30.7s
#17 Grok 4.20 Beta medium X AI 10.0 8.2 1/1 31.4s
#29 Hy3 preview medium Tencent 10.0 7.8 1/1 31.4s
#46 Claude Sonnet 4.6 medium Anthropic 6.9 7.6 1/2 33.9s
#74 Laguna M.1 medium Poolside 4.3 6.9 0/1 35.6s
#126 Kimi K2.5 none Moonshot AI 6.8 5.3 1/2 36.0s
#118 Nemotron 3 Nano Omni 30b A3b Reasoning medium NVIDIA 3.3 5.4 0/1 38.1s
#9 Gemini 3.5 Flash none Google 8.2 8.9 1/2 39.6s
#106 Owl Alpha none Openrouter 7.0 5.7 1/2 39.7s
#121 Mistral Small 4 medium Mistral 5.1 5.4 0/2 44.8s
#111 gpt-oss-120b medium OpenAI 3.9 5.6 0/2 47.2s
#94 GPT-5 Nano medium OpenAI 5.4 6.1 0/2 47.8s
#80 DeepSeek V4 Pro high DeepSeek 2.8 6.6 0/2 51.8s
#59 Qwen3.6 Flash medium Qwen 5.1 7.4 0/2 51.9s
#28 GLM 5 Turbo medium Z.ai 7.3 7.9 1/2 53.9s

Top Models by Coding Score

Coding Score vs Total Cost

Top Models by Response Time (avg)