Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #359 | MiniMax M2.5 medium | Minimax | 1 | 0.3 | $0.440 | 0/3 | 188.6s |
| #137 | Grok 4.7 low | X AI | 1 | 0.8 | $1.933 | 2/3 | 189.1s |
| #202 | Step 3.7 Flash high | Stepfun | 1 | 0.4 | $1.478 | 0/3 | 206.2s |
| #224 | Qwen3.5-35B-A3B medium | Qwen | 1 | 0.6 | $0.716 | 1/3 | 206.6s |
| #292 | Hy4 preview low | Tencent | 2 | 0.5 | $1.844 | 0/3 | 219.1s |
| #175 | Gemma 4 31B medium | 1 | 0.4 | $0.144 | 0/3 | 219.8s | |
| #129 | Seed-2.0-Mini medium | Bytedance Seed | 1 | 0.6 | $0.140 | 1/3 | 220.5s |
| #344 | Nemotron 3.5 Lightning low | NVIDIA | 3 | 0.4 | $0.175 | 0/3 | 224.4s |
| #92 | DeepSeek V4 Pro 0423 high | DeepSeek | 2 | 0.6 | $0.149 | 1/3 | 243.0s |
| #229 | Solar Pro 4 high | Upstage | 1 | 0.6 | $0.077 | 1/3 | 244.7s |
| #269 | Trinity Large Thinking high | Arcee AI | 1 | 0.4 | $0.794 | 0/3 | 245.0s |
| #234 | DeepSeek V3.2 medium | DeepSeek | 1 | 0.6 | $0.152 | 1/3 | 248.7s |
| #77 | DeepSeek V4 Flash 0731 high | DeepSeek | 1 | 0.6 | $0.576 | 1/3 | 252.7s |
| #189 | Solar Pro 4 low | Upstage | 1 | 0.5 | $0.101 | 1/3 | 259.6s |
| #131 | Seed-2.0-Code high | Bytedance Seed | 1 | 0.6 | $1.448 | 1/3 | 266.7s |