Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Tests Correct ↓.
Failure Reasons
197/197
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #214 | Qwen3.6 27B none | Qwen | 2 | 5.5 | $0.116 | 1/3 | 4.16s |
| #216 | Qwen3.8 27B none | Qwen | 2 | 5.5 | ~$0.006 | 1/3 | 1.19s |
| #219 | Laguna S 2.1 medium | Poolside | 1 | 5.3 | $0.053 | 1/3 | 159.0s |
| #223 | Kimi K2.5 none | Moonshot AI | 2 | 5.5 | $0.101 | 1/3 | 24.6s |
| #231 | Qwen3.6 35B A3B none | Qwen | 2 | 5.5 | $0.060 | 1/3 | 8.77s |
| #238 | MiMo-V2.5 none | Xiaomi | 2 | 5.5 | $0.025 | 1/3 | 3.24s |
| #128 | Qwen3.6 Flash medium | Qwen | 3 | 5.0 | $0.741 | 0/3 | 42.9s |
| #129 | Step 3.7 Flash high | Stepfun | 1 | 4.0 | $1.229 | 0/3 | 206.2s |
| #138 | GLM 5.2 none | Z.ai | 2 | 3.7 | $0.153 | 0/3 | 7.55s |
| #145 | Gemini 3.5 Flash Lite low | 3 | 4.3 | $0.138 | 0/3 | 917ms | |
| #151 | Qwen3.5 Plus 2026-02-15 none | Qwen | 3 | 4.3 | $0.073 | 0/3 | 2.05s |
| #156 | Qwen3.6 Max Preview none | Qwen | 3 | 3.8 | $0.228 | 0/3 | 3.12s |
| #168 | Claude Sonnet 5 none | Anthropic | 3 | 4.6 | $0.548 | 0/3 | 3.67s |
| #174 | Qwen3.5-Flash medium | Qwen | 2 | 3.7 | $0.142 | 0/3 | 58.9s |
| #175 | Inkling low | Thinkingmachines | 3 | 5.1 | $0.181 | 0/3 | 8.71s |