Puzzle Solving: API error
Puzzle Solving
API error
See which AI models are most likely to hit API error on Puzzle Solving, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
14/14
Filter models
No models match the current search and filters.
| Rank | Model | Company | API error Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #264 | Qwen3.6 Plus Preview medium | Qwen | 2 | 5.3 | $0.000 | 1/3 | 7.52s |
| #308 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 1 | 2.9 | $0.000 | 0/3 | 1.40s |
| #309 | Nemotron 3 Nano Omni 30b A3b Reasoning none | NVIDIA | 1 | 3.0 | $0.000 | 0/3 | 532ms |
| #311 | LFM2-24B-A2B none | Liquid | 1 | 3.8 | $0.001 | 0/3 | 1.78s |
| #306 | Laguna Xs.2 none | Poolside | 1 | 5.3 | $0.004 | 1/3 | 650ms |
| #287 | Laguna M.1 none | Poolside | 1 | 3.0 | $0.009 | 0/3 | 891ms |
| #294 | Laguna Xs.2 medium | Poolside | 1 | 5.3 | $0.015 | 1/3 | 1.93s |
| #279 | Laguna M.1 medium | Poolside | 1 | 5.3 | $0.033 | 1/3 | 10.2s |
| #268 | Hy4 preview none | Tencent | 1 | 4.6 | $0.041 | 0/3 | 24.1s |
| #230 | Hy3 preview low | Tencent | 1 | 5.3 | $0.042 | 1/3 | 7.51s |
| #262 | DeepSeek V3.2 none | DeepSeek | 1 | 7.6 | $0.054 | 2/3 | 6.91s |
| #205 | Hy3 preview high | Tencent | 1 | 7.7 | $0.135 | 2/3 | 27.9s |
| #138 | Seed-2.0-Code high | Bytedance Seed | 1 | 7.8 | $1.273 | 2/3 | 15.9s |
| #228 | Hy4 preview low | Tencent | 1 | 5.3 | $1.745 | 0/3 | 238.0s |