Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Tests Correct ↓.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #307 | Mimo V2 PRO medium | Xiaomi | 1 | 0.6 | $0.333 | 1/3 | 94.2s |
| #308 | MiMo-V2-Flash medium | Xiaomi | 1 | 0.6 | $0.043 | 1/3 | 10.7s |
| #311 | Mercury 2.5 low | Inception | 2 | 0.5 | $0.020 | 1/3 | 972ms |
| #315 | Laguna S 2.1 medium | Poolside | 1 | 0.5 | $0.092 | 1/3 | 159.0s |
| #316 | Mercury 2.5 Preview low | Inception | 2 | 0.5 | $0.015 | 1/3 | 972ms |
| #327 | MiMo-V2.5 none | Xiaomi | 2 | 0.5 | $0.049 | 1/3 | 3.24s |
| #342 | GLM 5V Turbo none | Z.ai | 2 | 0.5 | $0.052 | 1/3 | 3.13s |
| #345 | Owl Alpha medium | Openrouter | 1 | 0.5 | $0.000 | 1/3 | 18.7s |
| #346 | Mimo V2 PRO default | Xiaomi | 1 | 0.5 | $0.045 | 1/3 | 2.65s |
| #347 | Owl Alpha default | Openrouter | 1 | 0.6 | $0.000 | 1/3 | 36.9s |
| #355 | Granite 4.2 8B none | IBM Granite | 2 | 0.5 | $0.037 | 1/3 | 105.6s |
| #128 | GLM 5.1 medium | Z.ai | 1 | 0.5 | $1.546 | 0/3 | 109.6s |
| #133 | Qwen3.6 Flash medium | Qwen | 3 | 0.5 | $0.783 | 0/3 | 42.9s |
| #175 | Gemma 4 31B medium | 1 | 0.4 | $0.144 | 0/3 | 219.8s | |
| #180 | Qwen3.5-Flash medium | Qwen | 2 | 0.4 | $0.184 | 0/3 | 58.9s |