Combined Ranking
See which AI models perform best on Combined, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↑.
330/330
Filter models
No models match the current search and filters.
| Rank | Model | Company | Combined Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #268 | Mistral Small 4 medium | Mistral | 3.0 | 5.1 | $0.097 | 0/2 | 32.4s |
| #65 | Claude Fable 5.1 medium | Anthropic | 10.0 | 8.3 | $2.909 | 2/2 | 32.6s |
| #28 | GPT-5.5 medium | OpenAI | 10.0 | 9.0 | $4.163 | 2/2 | 33.5s |
| #26 | Claude Fable 5.1 high | Anthropic | 10.0 | 9.0 | $3.364 | 2/2 | 34.6s |
| #38 | Muse Spark 1.3 medium | Meta | 7.3 | 8.8 | $1.469 | 1/2 | 34.7s |
| #281 | Qwen3.6 Plus Preview medium | Qwen | 5.0 | 4.9 | $0.000 | 1/1 | 35.0s |
| #225 | Solar Pro 4 none | Upstage | 3.0 | 5.8 | $0.017 | 0/2 | 35.1s |
| #30 | Grok 4.5 high | X AI | 10.0 | 9.0 | $2.252 | 2/2 | 35.6s |
| #283 | Ling-2.6-flash none | Inclusionai | 3.0 | 4.9 | $0.002 | 0/2 | 35.7s |
| #312 | Hy3 preview none | Tencent | 1.5 | 4.0 | $0.007 | 0/1 | 35.8s |
| #91 | Claude Opus 4.8 low | Anthropic | 9.9 | 7.8 | $2.089 | 2/2 | 36.9s |
| #238 | Gemma 4 26B A4B none | 3.0 | 5.6 | $0.017 | 0/2 | 37.2s | |
| #114 | Qwen3.7 Max none | Qwen | 6.5 | 7.4 | $0.197 | 1/2 | 37.2s |
| #121 | Claude Sonnet 4.6 none | Anthropic | 9.8 | 7.3 | $0.661 | 2/2 | 37.5s |
| #295 | Grok 4.1 Fast medium | X AI | 5.0 | 4.7 | $0.069 | 1/1 | 37.6s |