Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.
404/404
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #241 | DeepSeek V4 Flash 0731 none | DeepSeek | 0.4 | 0.6 | $0.053 | 0/3 | 1.72s |
| #332 | GLM 4.7 Flash none | Z.ai | 0.4 | 0.5 | $0.027 | 0/3 | 2.54s |
| #306 | KAT Coder AIR V2.5 high | Kwaipilot | 0.4 | 0.5 | $0.077 | 0/3 | 39.1s |
| #301 | Laguna XS 2.1 none | Poolside | 0.4 | 0.5 | $0.020 | 0/3 | 623ms |
| #313 | MiMo-V2.5-Pro none | Xiaomi | 0.4 | 0.5 | $0.135 | 0/3 | 1.41s |
| #351 | Mimo V2 Omni default | Xiaomi | 0.4 | 0.4 | $0.021 | 0/3 | 2.75s |
| #305 | Nemotron 3.5 Lightning high | NVIDIA | 0.4 | 0.5 | $0.168 | 0/3 | 168.4s |
| #309 | Mistral Small 4 medium | Mistral | 0.4 | 0.5 | $0.126 | 0/3 | 40.0s |
| #334 | Nemotron 3.5 Lightning medium | NVIDIA | 0.4 | 0.5 | $0.195 | 0/3 | 285.5s |
| #247 | Inkling none | Thinkingmachines | 0.4 | 0.6 | $0.292 | 0/3 | 1.01s |
| #371 | Granite 4.1 8b default | IBM Granite | 0.4 | 0.4 | $0.007 | 0/3 | 775ms |
| #291 | North Mini Code medium | Cohere | 0.4 | 0.5 | $0.000 | 0/3 | 320.4s |
| #397 | Kev 4B default | Jaredpalmer | 0.5 | 0.2 | $0.002 | 1/2 | 771ms |
| #273 | Seed 2.1 Turbo none | Bytedance Seed | 0.5 | 0.6 | $0.182 | 0/3 | 2.35s |
| #283 | Ling 3.1 Flash none | Inclusionai | 0.5 | 0.5 | $0.000 | 0/3 | 31.9s |