Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↑.
289/289
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #129 | Step 3.7 Flash high | Stepfun | 4.0 | 6.8 | $1.229 | 0/3 | 206.2s |
| #170 | Qwen3.5-35B-A3B medium | Qwen | 5.9 | 6.1 | $1.140 | 1/3 | 206.6s |
| #283 | Ling 3.0 Tiny medium | Inclusionai | 2.9 | 3.8 | $0.000 | 0/3 | 212.0s |
| #105 | Kimi K2.6 medium | Moonshot AI | 5.7 | 7.2 | $1.235 | 1/3 | 214.4s |
| #123 | Kimi K2.5 medium | Moonshot AI | 6.1 | 7.0 | $0.479 | 1/3 | 217.5s |
| #166 | Gemma 4 31B medium | 4.3 | 6.2 | $0.100 | 0/3 | 219.8s | |
| #122 | Seed-2.0-Mini medium | Bytedance Seed | 5.5 | 7.0 | $0.116 | 1/3 | 220.5s |
| #250 | Nemotron 3.5 Lightning low | NVIDIA | 4.0 | 4.8 | $0.140 | 0/3 | 224.4s |
| #75 | DeepSeek V4 Pro high | DeepSeek | 6.3 | 7.7 | $0.761 | 1/3 | 243.0s |
| #114 | Solar Pro 4 high | Upstage | 5.7 | 7.1 | $0.046 | 1/3 | 244.7s |
| #252 | Trinity Large Thinking high | Arcee AI | 3.7 | 4.8 | $0.624 | 0/3 | 245.0s |
| #42 | Qwen3.8 27B high | Qwen | 8.1 | 8.4 | ~$0.287 | 2/3 | 248.6s |
| #117 | DeepSeek V3.2 medium | DeepSeek | 6.0 | 7.0 | $0.078 | 1/3 | 248.7s |
| #54 | DeepSeek V4 Flash 0731 high | DeepSeek | 6.3 | 8.0 | $0.134 | 1/3 | 252.7s |
| #189 | Step 3.5 Flash medium | Stepfun | 2.4 | 5.9 | $0.132 | 0/2 | 258.4s |