Puzzle Solving Ranking
See which AI models perform best on Puzzle Solving, which ones stay reliable, and where the biggest gaps appear. Sort by: Total Cost ↓.
311/311
Filter models
No models match the current search and filters.
| Rank | Model | Company | Puzzle Solving Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #277 | Hunter Alpha medium | OpenRouter | 6.1 | 4.7 | $0.000 | 1/3 | 5.35s |
| #280 | Cobuddy medium | Baidu | 3.6 | 4.7 | $0.000 | 0/3 | 12.8s |
| #288 | Elephant Alpha none | Openrouter | 4.2 | 4.3 | $0.000 | 0/3 | 807ms |
| #290 | Elephant Alpha medium | Openrouter | 5.3 | 4.3 | $0.000 | 1/3 | 868ms |
| #292 | Hunter Alpha none | OpenRouter | 5.8 | 4.2 | $0.000 | 1/3 | 3.71s |
| #297 | Ling 3.0 Tiny none | Inclusionai | 3.0 | 4.0 | $0.000 | 0/3 | 2.64s |
| #301 | Ling 3.0 Tiny high | Inclusionai | 3.0 | 3.8 | $0.000 | 0/3 | 54.2s |
| #302 | Ling 3.0 Tiny low | Inclusionai | 3.0 | 3.8 | $0.000 | 0/3 | 32.5s |
| #305 | Ling 3.0 Tiny medium | Inclusionai | 3.6 | 3.8 | $0.000 | 0/3 | 34.7s |
| #308 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 2.9 | 3.4 | $0.000 | 0/3 | 1.40s |
| #309 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 3.0 | 3.2 | $0.000 | 0/3 | 532ms |