Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #99 | Gemini 3.8 Flash medium | 1 | 0.8 | $0.769 | 2/3 | 11.7s | |
| #18 | GPT-6 Sol medium | OpenAI | 1 | 0.8 | $0.703 | 2/3 | 11.8s |
| #249 | Granite 4.2 8B medium | IBM Granite | 2 | 0.5 | $0.029 | 1/3 | 12.1s |
| #183 | GPT-6 Sol none | OpenAI | 2 | 0.5 | $0.551 | 1/3 | 12.2s |
| #363 | Mercury 2.5 none | Inception | 2 | 0.4 | $0.026 | 0/3 | 12.5s |
| #93 | Claude Opus 4.7 medium | Anthropic | 1 | 0.8 | $3.636 | 2/3 | 13.0s |
| #298 | Granite 4.2 8B low | IBM Granite | 2 | 0.5 | $0.036 | 1/3 | 13.0s |
| #139 | DeepSeek V4 Pro 0423 none | DeepSeek | 1 | 0.6 | $0.083 | 1/3 | 13.4s |
| #288 | Seed-2.0-Code none | Bytedance Seed | 3 | 0.4 | $0.257 | 0/3 | 13.4s |
| #29 | Claude Opus 5.5 medium | Anthropic | 1 | 0.8 | $1.804 | 2/3 | 13.5s |
| #104 | Claude Fable 5.1 high | Anthropic | 1 | 0.8 | $3.471 | 2/3 | 13.9s |
| #353 | KAT Coder AIR V2.5 default | Kwaipilot | 1 | 0.3 | $0.070 | 0/3 | 14.3s |
| #354 | Trinity Large Preview default | Arcee AI | 1 | 0.4 | $0.008 | 0/3 | 14.3s |
| #383 | Laguna Xs.2 medium | Poolside | 1 | 0.2 | $0.015 | 0/1 | 14.4s |
| #339 | DeepSeek V3.2 none | DeepSeek | 2 | 0.3 | $0.119 | 0/3 | 14.5s |