Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.
289/289
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #278 | Grok Build 0.1 none | X AI | 3.3 | 4.0 | $0.547 | 1/1 | 21.4s |
| #287 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 3.3 | 3.2 | $0.000 | 1/1 | 1.27s |
| #261 | Mercury 2 none | Inception | 3.4 | 4.7 | $0.030 | 0/3 | 1.03s |
| #264 | Laguna S 2.1 none | Poolside | 3.4 | 4.5 | $0.022 | 0/3 | 550ms |
| #263 | MiniMax M2.5 medium | Minimax | 3.4 | 4.6 | $0.273 | 0/3 | 188.6s |
| #220 | KAT-Coder-Air V2.5 low | Kwaipilot | 3.5 | 5.4 | $0.046 | 0/3 | 11.1s |
| #226 | Inkling Small none | Thinkingmachines | 3.6 | 5.3 | $0.054 | 0/3 | 759ms |
| #228 | Seed-2.0-Code none | Bytedance Seed | 3.6 | 5.3 | $0.151 | 0/3 | 13.4s |
| #212 | KAT-Coder-Air V2.5 medium | Kwaipilot | 3.6 | 5.6 | $0.048 | 0/3 | 23.4s |
| #256 | Trinity Large Preview none | Arcee AI | 3.7 | 4.8 | $0.008 | 0/3 | 14.3s |
| #204 | Qwen3.5-122B-A10B none | Qwen | 3.7 | 5.7 | $0.247 | 0/3 | 2.77s |
| #252 | Trinity Large Thinking high | Arcee AI | 3.7 | 4.8 | $0.624 | 0/3 | 245.0s |
| #138 | GLM 5.2 none | Z.ai | 3.7 | 6.7 | $0.153 | 0/3 | 7.55s |
| #235 | Mistral Small 4 none | Mistral | 3.7 | 5.1 | $0.022 | 0/3 | 901ms |
| #269 | Elephant Alpha medium | Openrouter | 3.7 | 4.3 | $0.000 | 0/3 | 1.30s |