Puzzle Solving Ranking
See which AI models perform best on Puzzle Solving, which ones stay reliable, and where the biggest gaps appear.
311/311
Filter models
No models match the current search and filters.
| Rank | Model | Company | Puzzle Solving Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #101 | Qwen3.7 Max none | Qwen | 10.0 | 7.4 | $0.197 | 3/3 | 1.13s |
| #105 | Qwen3.8 2.4T A95B high | Qwen | 10.0 | 7.4 | $2.357 | 3/3 | 65.5s |
| #106 | Grok 4.6 low | X AI | 10.0 | 7.4 | $0.449 | 3/3 | 8.16s |
| #107 | Gemini 3 Flash Preview low | 10.0 | 7.4 | $0.185 | 3/3 | 5.77s | |
| #111 | GLM 5.3 Flash high | Z.ai | 10.0 | 7.3 | $0.029 | 3/3 | 3.04s |
| #122 | Qwen3.5-122B-A10B medium | Qwen | 10.0 | 7.1 | $1.207 | 3/3 | 17.9s |
| #139 | Gemini 3.5 Flash none | 10.0 | 6.9 | $1.093 | 3/3 | 3.13s | |
| #145 | DeepSeek V4 Pro none | DeepSeek | 10.0 | 6.8 | $0.226 | 3/3 | 3.61s |
| #146 | GLM 5.3 low | Z.ai | 10.0 | 6.8 | $0.257 | 3/3 | 4.20s |
| #149 | Gemma 4 26B A4B medium | 10.0 | 6.8 | $0.092 | 3/3 | 5.79s | |
| #153 | Gemini 3.5 Flash minimal | 10.0 | 6.7 | $0.300 | 3/3 | 1.45s | |
| #156 | Claude Opus 4.7 none | Anthropic | 10.0 | 6.6 | $0.505 | 3/3 | 2.46s |
| #159 | Gemini 3.1 Flash Lite Preview low | 10.0 | 6.6 | $0.646 | 3/3 | 1.69s | |
| #161 | Gemini 3.1 Flash Lite low | 10.0 | 6.6 | $0.621 | 3/3 | 1.40s | |
| #163 | Mercury 2.5 Preview high | Inception | 10.0 | 6.6 | $0.030 | 3/3 | 1.21s |