Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Tests Correct ↓.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #261 | Claude Sonnet 5 none | Anthropic | 3 | 0.5 | $1.095 | 0/3 | 3.67s |
| #269 | Trinity Large Thinking high | Arcee AI | 1 | 0.4 | $0.794 | 0/3 | 245.0s |
| #272 | Mistral Large 4 none | Mistral | 3 | 0.5 | $0.290 | 0/3 | 85.0s |
| #273 | Seed 2.1 Turbo none | Bytedance Seed | 3 | 0.5 | $0.182 | 0/3 | 2.35s |
| #274 | Dots 3 Note Preview medium | Dots Studio | 1 | 0.4 | $0.000 | 0/3 | 309.1s |
| #279 | GLM 5 none | Z.ai | 3 | 0.4 | $0.140 | 0/3 | 5.12s |
| #283 | Ling 3.1 Flash none | Inclusionai | 1 | 0.5 | $0.000 | 0/3 | 31.9s |
| #286 | Gemma 4 26B A4B none | 2 | 0.4 | $0.031 | 0/3 | 4.16s | |
| #288 | Seed-2.0-Code none | Bytedance Seed | 3 | 0.4 | $0.257 | 0/3 | 13.4s |
| #291 | North Mini Code medium | Cohere | 3 | 0.4 | $0.000 | 0/3 | 320.4s |
| #292 | Hy4 preview low | Tencent | 2 | 0.5 | $1.844 | 0/3 | 219.1s |
| #293 | MiMo-V2.6-Flash none | Xiaomi | 2 | 0.4 | $0.060 | 0/3 | 744ms |
| #297 | GPT-5.6 Luna none | OpenAI | 3 | 0.4 | $0.042 | 0/3 | 980ms |
| #300 | Qwen3.6 Max Preview none | Qwen | 3 | 0.4 | $0.228 | 0/3 | 3.12s |
| #301 | Laguna XS 2.1 none | Poolside | 3 | 0.4 | $0.020 | 0/3 | 623ms |