Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Tests Correct ↑.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #293 | MiMo-V2.6-Flash none | Xiaomi | 2 | 0.4 | $0.060 | 0/3 | 744ms |
| #297 | GPT-5.6 Luna none | OpenAI | 3 | 0.4 | $0.042 | 0/3 | 980ms |
| #300 | Qwen3.6 Max Preview none | Qwen | 3 | 0.4 | $0.228 | 0/3 | 3.12s |
| #301 | Laguna XS 2.1 none | Poolside | 3 | 0.4 | $0.020 | 0/3 | 623ms |
| #303 | Inkling Small low | Thinkingmachines | 3 | 0.5 | $0.055 | 0/3 | 3.70s |
| #305 | Nemotron 3.5 Lightning high | NVIDIA | 2 | 0.4 | $0.168 | 0/3 | 168.4s |
| #306 | KAT Coder AIR V2.5 high | Kwaipilot | 1 | 0.4 | $0.077 | 0/3 | 39.1s |
| #309 | Mistral Small 4 medium | Mistral | 3 | 0.4 | $0.126 | 0/3 | 40.0s |
| #310 | Mistral Small 4 none | Mistral | 3 | 0.4 | $0.047 | 0/3 | 901ms |
| #312 | KAT Coder AIR V2.5 medium | Kwaipilot | 2 | 0.4 | $0.048 | 0/3 | 23.4s |
| #313 | MiMo-V2.5-Pro none | Xiaomi | 2 | 0.4 | $0.135 | 0/3 | 1.41s |
| #314 | Apodex 1.1 Mini default | Apodex | 3 | 0.4 | $0.000 | 0/3 | 1.13s |
| #318 | Dots 3 Note Preview default | Dots Studio | 2 | 0.5 | $0.000 | 0/3 | 6.12s |
| #319 | GPT-4o-mini default | OpenAI | 3 | 0.3 | $0.024 | 0/3 | 1.63s |
| #320 | KAT Coder AIR V2.5 low | Kwaipilot | 1 | 0.3 | $0.046 | 0/3 | 11.1s |