Combined Ranking
See which AI models perform best on Combined, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.
220/220
Filter models
No models match the current search and filters.
| Rank | Model | Company | Combined Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #169 | Qwen3.6 35B A3B none | Qwen | 3.8 | 5.3 | $0.061 | 0/2 | 39.5s |
| #97 | KAT-Coder-Pro V2.5 none | Kwaipilot | 4.1 | 6.7 | $0.476 | 0/2 | 183.1s |
| #156 | DeepSeek V4 Flash none | DeepSeek | 4.6 | 5.6 | $0.044 | 0/2 | 179.6s |
| #99 | Claude Opus 4.7 none | Anthropic | 4.8 | 6.6 | $0.505 | 1/1 | 18.3s |
| #182 | DeepSeek V3.2 none | DeepSeek | 4.8 | 5.0 | $0.054 | 0/2 | 113.5s |
| #119 | MiMo-V2-Flash medium | Xiaomi | 4.9 | 6.3 | $0.043 | 1/1 | 75.7s |
| #42 | GLM 5.2 medium | Z.ai | 5.0 | 7.8 | $0.182 | 1/1 | 52.0s |
| #46 | GLM 5 medium | Z.ai | 5.0 | 7.7 | $0.307 | 1/1 | 29.0s |
| #53 | GLM 5 Turbo medium | Z.ai | 5.0 | 7.6 | $0.323 | 1/1 | 13.9s |
| #106 | Hy3 preview medium | Tencent | 5.0 | 6.5 | $0.018 | 1/1 | 46.0s |
| #137 | Grok 4.20 Beta medium | X AI | 5.0 | 6.0 | $0.750 | 1/1 | 20.9s |
| #140 | Mimo V2 Omni medium | Xiaomi | 5.0 | 5.9 | $0.683 | 1/1 | 25.9s |
| #141 | Hy3 preview high | Tencent | 5.0 | 5.9 | $0.048 | 1/1 | 113.1s |
| #149 | Gemini 3.1 Flash Lite high | 5.0 | 5.6 | $2.044 | 1/1 | 149.2s | |
| #159 | Hy3 preview low | Tencent | 5.0 | 5.5 | $0.015 | 1/1 | 78.7s |