Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
197/197
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #156 | Qwen3.6 Max Preview none | Qwen | 3 | 3.8 | $0.228 | 0/3 | 3.12s |
| #207 | GLM 5V Turbo none | Z.ai | 2 | 5.5 | $0.052 | 1/3 | 3.13s |
| #238 | MiMo-V2.5 none | Xiaomi | 2 | 5.5 | $0.025 | 1/3 | 3.24s |
| #104 | Claude Opus 4.8 none | Anthropic | 1 | 5.5 | $1.166 | 1/3 | 3.29s |
| #168 | Claude Sonnet 5 none | Anthropic | 3 | 4.6 | $0.548 | 0/3 | 3.67s |
| #200 | Inkling Small low | Thinkingmachines | 3 | 4.6 | $0.055 | 0/3 | 3.70s |
| #101 | Gemini 3.1 Flash Lite medium | 2 | 5.5 | $0.120 | 1/3 | 3.81s | |
| #90 | Gemini 3.1 Flash Lite Preview medium | 2 | 5.5 | $0.117 | 1/3 | 4.09s | |
| #209 | Gemma 4 26B A4B none | 2 | 3.7 | $0.015 | 0/3 | 4.16s | |
| #214 | Qwen3.6 27B none | Qwen | 2 | 5.5 | $0.116 | 1/3 | 4.16s |
| #273 | Hy3 preview none | Tencent | 1 | 2.7 | $0.007 | 0/3 | 4.56s |
| #167 | GPT-5.6 Luna low | OpenAI | 2 | 5.5 | $0.051 | 1/3 | 4.61s |
| #202 | GLM 5.1 none | Z.ai | 3 | 3.9 | $0.164 | 0/3 | 4.96s |
| #199 | GLM 5 none | Z.ai | 3 | 4.0 | $0.027 | 0/3 | 5.12s |
| #98 | Claude Sonnet 4.6 none | Anthropic | 1 | 5.5 | $0.661 | 1/3 | 5.19s |