Combined Ranking
See which AI models perform best on Combined, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↓.
330/330
Filter models
No models match the current search and filters.
| Rank | Model | Company | Combined Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #61 | Seed 2.1 Turbo medium | Bytedance Seed | 10.0 | 8.4 | $1.951 | 2/2 | 400.3s |
| #204 | Trinity Large Thinking medium | Arcee AI | 2.9 | 6.1 | $0.756 | 0/2 | 382.7s |
| #24 | Seed 2.1 Turbo low | Bytedance Seed | 10.0 | 9.1 | $1.546 | 2/2 | 360.0s |
| #157 | Solar Pro 4 low | Upstage | 6.3 | 6.8 | $0.130 | 1/2 | 358.1s |
| #76 | GLM 5.2 high | Z.ai | 10.0 | 8.0 | $1.721 | 2/2 | 321.5s |
| #151 | Qwen3.7 Flash medium | Qwen | 6.4 | 6.9 | $0.065 | 1/2 | 314.2s |
| #135 | Qwen3.5-122B-A10B medium | Qwen | 6.4 | 7.1 | $1.047 | 1/2 | 313.5s |
| #154 | Qwen3.6 Flash medium | Qwen | 6.5 | 6.8 | $0.741 | 1/2 | 299.2s |
| #269 | Granite 4.2 8B low | IBM Granite | 3.0 | 5.1 | $0.012 | 0/2 | 298.2s |
| #158 | Solar Pro 4 medium | Upstage | 6.4 | 6.8 | $0.140 | 1/2 | 294.8s |
| #265 | Granite 4.2 8B medium | IBM Granite | 3.0 | 5.2 | $0.008 | 0/2 | 291.1s |
| #209 | Trinity Large Thinking low | Arcee AI | 5.2 | 6.0 | $0.625 | 0/2 | 291.0s |
| #25 | Qwen3.7 Max medium | Qwen | 8.7 | 9.1 | $1.157 | 1/2 | 287.8s |
| #250 | Laguna S 2.1 medium | Poolside | 3.2 | 5.4 | $0.053 | 0/2 | 284.7s |
| #147 | Seed-2.0-Mini medium | Bytedance Seed | 7.3 | 7.0 | $0.116 | 1/2 | 282.3s |