Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↓.
Failure Reasons
197/197
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #113 | Step 3.7 Flash low | Stepfun | 1 | 8.2 | $0.390 | 2/3 | 9.46s |
| #57 | GPT-5.6 Terra high | OpenAI | 1 | 7.6 | $0.836 | 2/3 | 9.14s |
| #7 | Gemini 3.7 Flash medium | 1 | 8.4 | $0.155 | 2/3 | 8.86s | |
| #231 | Qwen3.6 35B A3B none | Qwen | 2 | 5.5 | $0.060 | 1/3 | 8.77s |
| #175 | Inkling low | Thinkingmachines | 3 | 5.1 | $0.181 | 0/3 | 8.71s |
| #70 | Claude Opus 4.8 low | Anthropic | 1 | 6.6 | $2.089 | 1/3 | 7.58s |
| #138 | GLM 5.2 none | Z.ai | 2 | 3.7 | $0.153 | 0/3 | 7.55s |
| #66 | GPT-5.6 Terra medium | OpenAI | 2 | 6.1 | $0.523 | 1/3 | 7.19s |
| #28 | Gemini 3.6 Flash low | 1 | 7.8 | $0.251 | 2/3 | 6.95s | |
| #24 | Gemini 3.5 Flash low | 1 | 7.8 | $0.449 | 2/3 | 6.71s | |
| #27 | Claude Opus 5 low | Anthropic | 1 | 7.8 | $1.121 | 2/3 | 6.70s |
| #201 | Ling-3.0-flash none | Inclusionai | 2 | 5.5 | $0.004 | 1/3 | 6.39s |
| #13 | Gemini 3.7 Flash low | 1 | 7.6 | $0.104 | 2/3 | 6.20s | |
| #232 | Dots 3 Note Preview none | Dots Studio | 2 | 4.6 | $0.000 | 0/3 | 6.12s |
| #97 | Gemini 3 Flash Preview low | 2 | 5.8 | $0.185 | 1/3 | 6.00s |