Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↑.
289/289
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #118 | GLM 5.1 medium | Z.ai | 4.6 | 7.0 | $0.727 | 0/3 | 109.6s |
| #128 | Qwen3.6 Flash medium | Qwen | 5.0 | 6.8 | $0.741 | 0/3 | 42.9s |
| #129 | Step 3.7 Flash high | Stepfun | 4.0 | 6.8 | $1.229 | 0/3 | 206.2s |
| #137 | Gemma 4 26B A4B medium | 2.9 | 6.8 | $0.092 | 0/3 | 272.5s | |
| #138 | GLM 5.2 none | Z.ai | 3.7 | 6.7 | $0.153 | 0/3 | 7.55s |
| #145 | Gemini 3.5 Flash Lite low | 4.3 | 6.6 | $0.138 | 0/3 | 917ms | |
| #151 | Qwen3.5 Plus 2026-02-15 none | Qwen | 4.3 | 6.6 | $0.073 | 0/3 | 2.05s |
| #156 | Qwen3.6 Max Preview none | Qwen | 3.8 | 6.4 | $0.228 | 0/3 | 3.12s |
| #166 | Gemma 4 31B medium | 4.3 | 6.2 | $0.100 | 0/3 | 219.8s | |
| #168 | Claude Sonnet 5 none | Anthropic | 4.6 | 6.1 | $0.548 | 0/3 | 3.67s |
| #174 | Qwen3.5-Flash medium | Qwen | 3.7 | 6.1 | $0.142 | 0/3 | 58.9s |
| #175 | Inkling low | Thinkingmachines | 5.1 | 6.1 | $0.181 | 0/3 | 8.71s |
| #179 | Qwen3.5 Plus 2026-04-20 none | Qwen | 3.9 | 6.1 | $0.122 | 0/3 | 1.69s |
| #185 | Gemini 3 PRO Preview medium | 3.0 | 6.0 | $0.385 | 0/3 | 0ms | |
| #189 | Step 3.5 Flash medium | Stepfun | 2.4 | 5.9 | $0.132 | 0/2 | 258.4s |