Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Tests Correct ↑.
Failure Reasons
197/197
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #271 | Grok 4.20 none | X AI | 1 | 1.1 | $0.057 | 0/1 | 1.22s |
| #272 | Laguna Xs.2 medium | Poolside | 1 | 2.1 | $0.015 | 0/1 | 14.4s |
| #273 | Hy3 preview none | Tencent | 1 | 2.7 | $0.007 | 0/3 | 4.56s |
| #274 | MiMo-V2-Flash none | Xiaomi | 2 | 4.3 | $0.025 | 0/3 | 2.64s |
| #275 | Ling 3.0 Tiny none | Inclusionai | 3 | 4.2 | $0.000 | 0/3 | 1.19s |
| #276 | Granite 4.1 8B none | IBM Granite | 1 | 4.5 | $0.007 | 0/3 | 775ms |
| #277 | Nemotron 3.5 Lightning none | NVIDIA | 3 | 3.1 | $0.007 | 0/3 | 1.55s |
| #281 | Grok 4.1 Fast none | X AI | 1 | 1.8 | $0.008 | 0/1 | 1.79s |
| #284 | Laguna Xs.2 none | Poolside | 1 | 8.3 | $0.004 | 0/1 | 1.96s |
| #285 | gpt-oss-120b none | OpenAI | 1 | 1.5 | $0.010 | 0/1 | 9.57s |
| #286 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 1 | 1.1 | $0.000 | 0/1 | 38.1s |
| #53 | DeepSeek V4 Flash 0731 low | DeepSeek | 1 | 7.3 | $0.072 | 1/3 | 163.3s |
| #62 | Qwen3.7 Plus medium | Qwen | 1 | 6.1 | $0.277 | 1/3 | 108.6s |
| #63 | Qwen3.6 Plus medium | Qwen | 1 | 6.1 | $0.418 | 1/3 | 153.1s |
| #66 | GPT-5.6 Terra medium | OpenAI | 2 | 6.1 | $0.523 | 1/3 | 7.19s |