Puzzle Solving Ranking
See which AI models perform best on Puzzle Solving, which ones stay reliable, and where the biggest gaps appear.
311/311
Filter models
No models match the current search and filters.
| Rank | Model | Company | Puzzle Solving Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #204 | Mimo V2 Omni medium | Xiaomi | 5.9 | 5.9 | $0.683 | 1/3 | 2.38s |
| #259 | MiniMax M2.7 medium | Minimax | 5.9 | 5.0 | $0.208 | 1/3 | 24.9s |
| #102 | Kimi K2.7 Code medium | Moonshot AI | 5.9 | 7.4 | $0.687 | 1/3 | 41.0s |
| #160 | Inkling Small medium | Thinkingmachines | 5.8 | 6.6 | $0.113 | 1/3 | 5.24s |
| #292 | Hunter Alpha none | OpenRouter | 5.8 | 4.2 | $0.000 | 1/3 | 3.71s |
| #219 | Gemini 3.1 Flash Lite high | 5.7 | 5.6 | $2.044 | 1/3 | 50.8s | |
| #68 | Step 3.7 Flash medium | Stepfun | 5.7 | 7.9 | $0.495 | 1/3 | 6.19s |
| #57 | GPT-5 Mini medium | OpenAI | 5.6 | 8.1 | $0.241 | 1/3 | 15.2s |
| #248 | Inkling none | Thinkingmachines | 5.6 | 5.2 | $0.147 | 1/3 | 931ms |
| #209 | GPT-5.4 none | OpenAI | 5.6 | 5.8 | $0.397 | 1/3 | 1.44s |
| #265 | Mercury 2.5 Preview none | Inception | 5.6 | 4.9 | $0.018 | 1/3 | 684ms |
| #171 | Dots 3 Note Preview high | Dots Studio | 5.5 | 6.4 | $0.000 | 1/3 | 12.7s |
| #124 | Step 3.7 Flash low | Stepfun | 5.5 | 7.1 | $0.390 | 1/3 | 1.84s |
| #267 | Nemotron 3 Super none | NVIDIA | 5.5 | 4.8 | $0.008 | 1/3 | 2.36s |
| #257 | GLM 5 Turbo none | Z.ai | 5.5 | 5.1 | $0.047 | 1/3 | 2.65s |