Combined Ranking
See which AI models perform best on Combined, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↓.
330/330
Filter models
No models match the current search and filters.
| Rank | Model | Company | Combined Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #261 | Gemini 3.1 Flash Lite Preview high | 5.0 | 5.3 | $2.310 | 1/1 | 280.5s | |
| #71 | Muse Glimmer 30B xhigh | Meta | 6.0 | 8.1 | $0.471 | 0/2 | 272.7s |
| #202 | Qwen3.5-Flash medium | Qwen | 6.4 | 6.1 | $0.142 | 1/2 | 266.6s |
| #118 | Qwen3.8 2.4T A95B high | Qwen | 6.9 | 7.4 | $2.357 | 1/2 | 266.1s |
| #289 | Trinity Large Thinking high | Arcee AI | 3.0 | 4.8 | $0.592 | 0/2 | 262.4s |
| #227 | Nemotron 3 Super medium | NVIDIA | 6.4 | 5.7 | $0.054 | 1/2 | 259.9s |
| #190 | Ring-2.6-1T medium | Inclusionai | 7.3 | 6.4 | $0.102 | 1/2 | 257.3s |
| #215 | Qwen3.5-Flash none | Qwen | 2.9 | 6.0 | $0.073 | 0/2 | 243.6s |
| #92 | Qwen3.7 Flash high | Qwen | 6.9 | 7.8 | $0.052 | 1/2 | 232.0s |
| #69 | Qwen3.8 2.4T A95B low | Qwen | 8.2 | 8.1 | $2.672 | 1/2 | 231.8s |
| #79 | Kimi K3 max | Moonshot AI | 6.5 | 7.9 | $2.528 | 1/2 | 223.0s |
| #101 | Nemotron 3 Ultra medium | NVIDIA | 6.3 | 7.6 | $0.701 | 1/2 | 218.2s |
| #181 | Laguna XS 2.1 medium | Poolside | 6.3 | 6.5 | $0.068 | 1/2 | 218.1s |
| #133 | Qwen3.7 Flash low | Qwen | 6.4 | 7.2 | $0.043 | 1/2 | 217.8s |
| #100 | Solar Pro 4 xhigh | Upstage | 10.0 | 7.6 | $0.150 | 2/2 | 202.5s |