Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear.
404/404
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #310 | Mistral Small 4 none | Mistral | 0.4 | 0.5 | $0.047 | 0/3 | 901ms |
| #269 | Trinity Large Thinking high | Arcee AI | 0.4 | 0.6 | $0.794 | 0/3 | 245.0s |
| #213 | Qwen3.5-122B-A10B none | Qwen | 0.4 | 0.6 | $0.308 | 0/3 | 2.77s |
| #363 | Mercury 2.5 none | Inception | 0.4 | 0.4 | $0.026 | 0/3 | 12.5s |
| #354 | Trinity Large Preview default | Arcee AI | 0.4 | 0.4 | $0.008 | 0/3 | 14.3s |
| #312 | KAT Coder AIR V2.5 medium | Kwaipilot | 0.4 | 0.5 | $0.048 | 0/3 | 23.4s |
| #288 | Seed-2.0-Code none | Bytedance Seed | 0.4 | 0.5 | $0.257 | 0/3 | 13.4s |
| #322 | Inkling Small none | Thinkingmachines | 0.4 | 0.5 | $0.054 | 0/3 | 759ms |
| #320 | KAT Coder AIR V2.5 low | Kwaipilot | 0.3 | 0.5 | $0.046 | 0/3 | 11.1s |
| #359 | MiniMax M2.5 medium | Minimax | 0.3 | 0.4 | $0.440 | 0/3 | 188.6s |
| #243 | Claude Fable 5.1 low | Anthropic | 0.3 | 0.6 | $3.395 | 0/3 | 6.75s |
| #204 | Claude Sonnet 5.5 medium | Anthropic | 0.3 | 0.7 | $1.100 | 0/3 | 6.07s |
| #336 | Laguna S 2.1 none | Poolside | 0.3 | 0.5 | $0.044 | 0/3 | 550ms |
| #357 | Mercury 2 none | Inception | 0.3 | 0.4 | $0.033 | 0/3 | 1.03s |
| #130 | Mistral Large 4 high | Mistral | 0.3 | 0.8 | $1.391 | 0/3 | 1277.8s |