Coding Ranking
See which AI models perform best on Coding, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.
404/404
Filter models
No models match the current search and filters.
| Rank | Model | Company | Coding Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #26 | Claude Opus 5 high | Anthropic | 1.0 | 0.9 | $4.524 | 3/3 | 12.7s |
| #27 | Claude Opus 5 medium | Anthropic | 1.0 | 0.9 | $2.870 | 3/3 | 9.56s |
| #28 | Seed 2.1 Turbo low | Bytedance Seed | 1.0 | 0.9 | $1.645 | 3/3 | 360.8s |
| #31 | Claude Sonnet 5.5 max | Anthropic | 1.0 | 0.9 | $8.557 | 3/3 | 55.3s |
| #32 | Qwen3.8 MAX medium | Qwen | 1.0 | 0.9 | $1.181 | 3/3 | 41.3s |
| #33 | Qwen3.8 Max (0902) high | Qwen | 1.0 | 0.9 | $3.306 | 3/3 | 253.2s |
| #34 | GPT-5.3-Codex medium | OpenAI | 1.0 | 0.9 | $1.318 | 3/3 | 19.5s |
| #35 | Claude Haiku 5.5 xhigh | Anthropic | 1.0 | 0.9 | $0.140 | 3/3 | 13.1s |
| #36 | Muse Spark 1.3 high | Meta | 1.0 | 0.9 | $1.891 | 3/3 | 51.7s |
| #37 | Pareto default | Unbiased | 1.0 | 0.9 | $1.520 | 3/3 | 50.1s |
| #38 | Seed 2.1 Turbo high | Bytedance Seed | 1.0 | 0.9 | $1.892 | 3/3 | 336.0s |
| #41 | Muse Spark 1.3 medium | Meta | 1.0 | 0.9 | $1.712 | 3/3 | 41.1s |
| #45 | GLM 5.3 Flash max | Z.ai | 1.0 | 0.9 | $0.089 | 3/3 | 43.4s |
| #46 | Muse Spark 1.2 high | Meta | 1.0 | 0.9 | $2.018 | 3/3 | 22.2s |
| #47 | Muse Spark 1.2 medium | Meta | 1.0 | 0.9 | $1.590 | 3/3 | 17.5s |