Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster.
Failure Reasons
218/218
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #153 | Qwen3.6 Flash medium | Qwen | 3 | 5.0 | $0.741 | 0/3 | 42.9s |
| #161 | Mercury 2.5 medium | Inception | 3 | 5.0 | $0.022 | 0/3 | 3.70s |
| #171 | Gemini 3.5 Flash Lite low | 3 | 4.3 | $0.138 | 0/3 | 917ms | |
| #178 | Qwen3.5 Plus 2026-02-15 none | Qwen | 3 | 4.3 | $0.073 | 0/3 | 2.05s |
| #183 | Qwen3.6 Max Preview none | Qwen | 3 | 3.8 | $0.228 | 0/3 | 3.12s |
| #195 | Claude Sonnet 5 none | Anthropic | 3 | 4.6 | $0.548 | 0/3 | 3.67s |
| #202 | Inkling low | Thinkingmachines | 3 | 5.1 | $0.187 | 0/3 | 8.71s |
| #222 | Seed 2.1 Turbo none | Bytedance Seed | 3 | 4.5 | $0.092 | 0/3 | 2.35s |
| #225 | North Mini Code medium | Cohere | 3 | 4.5 | $0.000 | 0/3 | 320.4s |
| #227 | GLM 5 none | Z.ai | 3 | 4.0 | $0.027 | 0/3 | 5.12s |
| #228 | Inkling Small low | Thinkingmachines | 3 | 4.6 | $0.055 | 0/3 | 3.70s |
| #230 | GLM 5.1 none | Z.ai | 3 | 3.9 | $0.164 | 0/3 | 4.96s |
| #232 | Qwen3.5-122B-A10B none | Qwen | 3 | 3.7 | $0.247 | 0/3 | 2.77s |
| #247 | DeepSeek V4 Flash 0423 none | DeepSeek | 3 | 4.2 | $0.040 | 0/3 | 17.1s |
| #251 | GPT-5.6 Luna none | OpenAI | 3 | 3.8 | $0.029 | 0/3 | 980ms |