Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Total Cost ↓.
289/289
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #42 | Qwen3.8 27B high | Qwen | 8.1 | 8.4 | ~$0.287 | 2/3 | 248.6s |
| #110 | Muse Glimmer 30B medium | Meta | 8.4 | 7.2 | $0.280 | 2/3 | 33.7s |
| #184 | GPT-5.6 Terra none | OpenAI | 5.5 | 6.0 | $0.280 | 1/3 | 1.00s |
| #62 | Qwen3.7 Plus medium | Qwen | 6.1 | 7.9 | $0.277 | 1/3 | 108.6s |
| #263 | MiniMax M2.5 medium | Minimax | 3.4 | 4.6 | $0.273 | 0/3 | 188.6s |
| #64 | Gemini 3.5 Flash Lite medium | 7.9 | 7.8 | $0.272 | 2/3 | 9.96s | |
| #116 | GPT-5.6 Sol none | OpenAI | 5.5 | 7.0 | $0.262 | 1/3 | 1.39s |
| #134 | GLM 5.3 low | Z.ai | 6.2 | 6.8 | $0.257 | 1/3 | 21.0s |
| #28 | Gemini 3.6 Flash low | 7.8 | 8.7 | $0.251 | 2/3 | 6.95s | |
| #204 | Qwen3.5-122B-A10B none | Qwen | 3.7 | 5.7 | $0.247 | 0/3 | 2.77s |
| #49 | GPT-5 Mini medium | OpenAI | 10.0 | 8.1 | $0.241 | 3/3 | 27.6s |
| #69 | Seed-2.0-Lite medium | Bytedance Seed | 8.0 | 7.8 | $0.236 | 2/3 | 156.7s |
| #203 | Kimi K2.6 none | Moonshot AI | 5.5 | 5.7 | $0.230 | 1/3 | 82.6s |
| #156 | Qwen3.6 Max Preview none | Qwen | 3.8 | 6.4 | $0.228 | 0/3 | 3.12s |
| #72 | GLM 5 medium | Z.ai | 10.0 | 7.7 | $0.228 | 3/3 | 74.3s |