Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Total Cost ↓.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #265 | GPT-5.6 Terra none | OpenAI | 2 | 0.5 | $0.377 | 1/3 | 1.00s |
| #44 | Gemini 3.6 Flash low | 1 | 0.8 | $0.363 | 2/3 | 6.95s | |
| #307 | Mimo V2 PRO medium | Xiaomi | 1 | 0.6 | $0.333 | 1/3 | 94.2s |
| #152 | Gemini 3.8 Flash low | 2 | 0.5 | $0.323 | 1/3 | 4.31s | |
| #213 | Qwen3.5-122B-A10B none | Qwen | 3 | 0.4 | $0.308 | 0/3 | 2.77s |
| #134 | DeepSeek V4 Flash 0731 low | DeepSeek | 1 | 0.7 | $0.306 | 1/3 | 163.3s |
| #181 | Inkling low | Thinkingmachines | 3 | 0.5 | $0.293 | 0/3 | 8.71s |
| #247 | Inkling none | Thinkingmachines | 3 | 0.4 | $0.292 | 0/3 | 1.01s |
| #272 | Mistral Large 4 none | Mistral | 3 | 0.5 | $0.290 | 0/3 | 85.0s |
| #123 | Muse Glimmer 30B medium | Meta | 1 | 0.8 | $0.288 | 2/3 | 33.7s |
| #17 | Gemini 3.7 Flash medium | 1 | 0.8 | $0.278 | 2/3 | 8.86s | |
| #90 | Seed-2.0-Lite medium | Bytedance Seed | 1 | 0.8 | $0.271 | 2/3 | 156.7s |
| #299 | Kimi K2.5 none | Moonshot AI | 2 | 0.5 | $0.262 | 1/3 | 24.6s |
| #94 | DeepSeek V4 Flash 0423 high | DeepSeek | 1 | 0.8 | $0.261 | 2/3 | 50.6s |
| #288 | Seed-2.0-Code none | Bytedance Seed | 3 | 0.4 | $0.257 | 0/3 | 13.4s |