Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↓.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #208 | Mercury 2.5 high | Inception | 2 | 0.6 | $0.044 | 1/3 | 6.58s |
| #307 | Ling 3.0 Flash none | Inclusionai | 2 | 0.5 | $0.004 | 1/3 | 6.39s |
| #161 | GLM 5.3 FlashX medium | Z.ai | 2 | 0.6 | $0.114 | 1/3 | 6.37s |
| #25 | Gemini 3.7 Flash low | 1 | 0.8 | $0.183 | 2/3 | 6.20s | |
| #207 | Mercury 2.5 Preview high | Inception | 1 | 0.7 | $0.040 | 1/3 | 6.13s |
| #321 | Dots 3 Note Preview default | Dots Studio | 2 | 0.5 | $0.000 | 0/3 | 6.12s |
| #172 | Gemini 3 Flash Preview low | 2 | 0.6 | $0.254 | 1/3 | 6.00s | |
| #331 | Qwen3.5-9B none | Qwen | 3 | 0.4 | $0.041 | 0/3 | 5.60s |
| #159 | Claude Opus 5 none | Anthropic | 1 | 0.6 | $2.766 | 1/3 | 5.54s |
| #255 | GLM 5.3 FlashX low | Z.ai | 2 | 0.5 | $0.157 | 1/3 | 5.45s |
| #254 | Claude Sonnet 4.6 none | Anthropic | 1 | 0.5 | $0.661 | 1/3 | 5.19s |
| #282 | GLM 5 none | Z.ai | 3 | 0.4 | $0.140 | 0/3 | 5.12s |
| #210 | GLM 5.1 none | Z.ai | 3 | 0.4 | $0.451 | 0/3 | 4.96s |
| #177 | GPT-5.6 Luna low | OpenAI | 2 | 0.5 | $0.049 | 1/3 | 4.61s |
| #387 | Hy3 preview none | Tencent | 1 | 0.3 | $0.007 | 0/3 | 4.56s |