Combined Ranking
See which AI models perform best on Combined, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.
311/311
Filter models
No models match the current search and filters.
| Rank | Model | Company | Combined Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #83 | DeepSeek V4 Pro high | DeepSeek | 10.0 | 7.7 | $0.489 | 2/2 | 79.0s |
| #86 | GPT-5.6 Luna high | OpenAI | 10.0 | 7.6 | $0.179 | 2/2 | 19.0s |
| #87 | Solar Pro 4 xhigh | Upstage | 10.0 | 7.6 | $0.050 | 2/2 | 202.5s |
| #92 | MiniMax M3 medium | Minimax | 10.0 | 7.5 | $0.300 | 2/2 | 138.2s |
| #93 | GPT 5.3 Chat none | OpenAI | 10.0 | 7.5 | $0.533 | 2/2 | 15.1s |
| #98 | Grok Build 0.1 medium | X AI | 10.0 | 7.5 | $1.130 | 2/2 | 65.1s |
| #103 | GPT-5.6 Luna medium | OpenAI | 10.0 | 7.4 | $0.072 | 2/2 | 14.6s |
| #113 | Qwen3.7 Plus none | Qwen | 10.0 | 7.3 | $0.106 | 2/2 | 117.7s |
| #136 | Claude Fable 5.1 low | Anthropic | 10.0 | 6.9 | $2.565 | 2/2 | 29.4s |
| #152 | LongCat 2.0 low | Meituan | 10.0 | 6.7 | $0.412 | 2/2 | 130.2s |
| #155 | LongCat 2.0 high | Meituan | 10.0 | 6.7 | $0.492 | 2/2 | 167.1s |