Coding: Wrong answer
Coding
Wrong answer
See which AI models are most likely to hit Wrong answer on Coding, so you can spot weak points faster. Sort by: Total Cost ↓.
Failure Reasons
265/265
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #255 | Inkling Small medium | Thinkingmachines | 1 | 0.8 | $0.113 | 2/3 | 10.5s |
| #337 | Granite 4.2 8B high | IBM Granite | 1 | 0.4 | $0.111 | 0/3 | 515.3s |
| #343 | Laguna S 2.1 low | Poolside | 3 | 0.4 | $0.108 | 0/3 | 167.5s |
| #302 | Qwen3.6 35B A3B none | Qwen | 2 | 0.5 | $0.105 | 1/3 | 8.77s |
| #223 | LongCat 2.0 none | Meituan | 2 | 0.5 | $0.104 | 1/3 | 2.85s |
| #189 | Solar Pro 4 low | Upstage | 1 | 0.5 | $0.101 | 1/3 | 259.6s |
| #215 | Solar Pro 4 medium | Upstage | 1 | 0.5 | $0.098 | 1/3 | 310.4s |
| #315 | Laguna S 2.1 medium | Poolside | 1 | 0.5 | $0.092 | 1/3 | 159.0s |
| #242 | Gemini 3.1 Flash Lite Preview none | 2 | 0.5 | $0.090 | 1/3 | 967ms | |
| #188 | Qwen3.5-Flash none | Qwen | 2 | 0.5 | $0.089 | 1/3 | 850ms |
| #80 | Claude Haiku 5.5 high | Anthropic | 2 | 0.6 | $0.088 | 1/3 | 9.54s |
| #376 | Grok 4.20 Beta default | X AI | 1 | 0.2 | $0.087 | 0/1 | 1.14s |
| #212 | Laguna XS 2.1 medium | Poolside | 2 | 0.5 | $0.083 | 1/3 | 70.3s |
| #139 | DeepSeek V4 Pro 0423 none | DeepSeek | 1 | 0.6 | $0.083 | 1/3 | 13.4s |
| #257 | Gemini 3.1 Flash Lite minimal | 2 | 0.5 | $0.080 | 1/3 | 831ms |