Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear.
408/408
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #193 | Ember-1 max | Fireworks | 0.5 | 0.7 | $3.264 | 1/3 | 105.1s |
| #182 | Inkling low | Thinkingmachines | 0.5 | 0.7 | $0.293 | 0/3 | 8.71s |
| #247 | Apodex 1.1 Mini low | Apodex | 0.5 | 0.6 | $0.000 | 0/3 | 48.8s |
| #133 | Qwen3.6 Flash medium | Qwen | 0.5 | 0.7 | $0.783 | 0/3 | 42.9s |
| #219 | Mercury 2.5 medium | Inception | 0.5 | 0.6 | $0.034 | 0/3 | 3.70s |
| #398 | Decider V1.1 27B default | Perplexity | 0.5 | 0.3 | $0.002 | 1/2 | 305ms |
| #274 | Mistral Large 4 none | Mistral | 0.5 | 0.6 | $0.290 | 0/3 | 85.0s |
| #336 | Mercury 2.5 Preview default | Inception | 0.5 | 0.5 | $0.023 | 0/3 | 459ms |
| #128 | GLM 5.1 medium | Z.ai | 0.5 | 0.8 | $1.546 | 0/3 | 109.6s |
| #343 | GPT-5.4 Nano none | OpenAI | 0.5 | 0.5 | $0.068 | 0/3 | 2.22s |
| #306 | Inkling Small low | Thinkingmachines | 0.5 | 0.5 | $0.055 | 0/3 | 3.70s |
| #263 | Claude Sonnet 5 none | Anthropic | 0.5 | 0.6 | $1.095 | 0/3 | 3.67s |
| #321 | Dots 3 Note Preview default | Dots Studio | 0.5 | 0.5 | $0.000 | 0/3 | 6.12s |
| #326 | Qwen3 Coder Next default | Qwen | 0.5 | 0.5 | $0.067 | 0/3 | 2.22s |
| #275 | Seed 2.1 Turbo none | Bytedance Seed | 0.5 | 0.6 | $0.182 | 0/3 | 2.35s |