Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #267 | Ling 3.0 Flash high | Inclusionai | 2 | 0.7 | $0.021 | 1/3 | 14.9s |
| #125 | GPT-6 Luna low | OpenAI | 2 | 0.6 | $0.033 | 1/3 | 15.4s |
| #96 | GPT-5.6 Luna high | OpenAI | 2 | 0.5 | $0.118 | 1/3 | 15.6s |
| #172 | Muse Glimmer 30B low | Meta | 1 | 0.8 | $0.219 | 2/3 | 16.5s |
| #42 | Gemini 3.5 Flash Lite high | 1 | 0.8 | $0.624 | 2/3 | 16.6s | |
| #184 | Space Bunny Alpha max | Stealth | 2 | 0.6 | $0.000 | 1/3 | 17.1s |
| #236 | DeepSeek V4 Flash 0423 none | DeepSeek | 3 | 0.4 | $0.149 | 0/3 | 17.1s |
| #68 | Claude Sonnet 5 medium | Anthropic | 1 | 0.9 | $1.438 | 2/3 | 17.3s |
| #109 | MiMo-V2.6-Flash medium | Xiaomi | 1 | 0.7 | $0.218 | 1/3 | 18.2s |
| #345 | Owl Alpha medium | Openrouter | 1 | 0.5 | $0.000 | 1/3 | 18.7s |
| #112 | GLM 5.3 high | Z.ai | 2 | 0.7 | $0.495 | 1/3 | 18.9s |
| #98 | GPT-5.4 Nano medium | OpenAI | 2 | 0.6 | $0.235 | 1/3 | 19.1s |
| #142 | GLM 5.3 Flash high | Z.ai | 1 | 0.8 | $0.056 | 2/3 | 19.6s |
| #208 | GLM 5.3 low | Z.ai | 2 | 0.6 | $0.215 | 1/3 | 21.0s |
| #170 | Grok 4.6 low | X AI | 2 | 0.6 | $0.781 | 1/3 | 21.7s |