Puzzle Solving Ranking
See which AI models perform best on Puzzle Solving, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↑.
311/311
Filter models
No models match the current search and filters.
| Rank | Model | Company | Puzzle Solving Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #158 | Qwen3.5-27B none | Qwen | 6.7 | 6.6 | $0.058 | 1/3 | 1.38s |
| #185 | Qwen3.7 Flash none | Qwen | 3.7 | 6.1 | $0.019 | 0/3 | 1.39s |
| #308 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 2.9 | 3.4 | $0.000 | 0/3 | 1.40s |
| #217 | Kimi K2.6 none | Moonshot AI | 3.1 | 5.7 | $0.233 | 0/3 | 1.40s |
| #161 | Gemini 3.1 Flash Lite low | 10.0 | 6.6 | $0.621 | 3/3 | 1.40s | |
| #209 | GPT-5.4 none | OpenAI | 5.6 | 5.8 | $0.397 | 1/3 | 1.44s |
| #153 | Gemini 3.5 Flash minimal | 10.0 | 6.7 | $0.300 | 3/3 | 1.45s | |
| #216 | GLM 5.1 none | Z.ai | 7.7 | 5.7 | $0.164 | 2/3 | 1.45s |
| #215 | Ling-3.0-flash none | Inclusionai | 5.3 | 5.7 | $0.004 | 1/3 | 1.49s |
| #127 | GPT-5.6 Sol none | OpenAI | 7.7 | 7.0 | $0.201 | 2/3 | 1.49s |
| #247 | Dots 3 Note Preview none | Dots Studio | 3.3 | 5.2 | $0.000 | 0/3 | 1.50s |
| #214 | Inkling Small low | Thinkingmachines | 5.3 | 5.7 | $0.055 | 1/3 | 1.56s |
| #235 | KAT-Coder-Air V2.5 low | Kwaipilot | 3.1 | 5.4 | $0.046 | 0/3 | 1.57s |
| #224 | Mimo V2 PRO none | Xiaomi | 6.0 | 5.6 | $0.045 | 1/3 | 1.61s |
| #240 | Laguna S 2.1 high | Poolside | 2.9 | 5.4 | $0.114 | 0/3 | 1.62s |