Combined: Wrong answer
Combined
Wrong answer
See which AI models are most likely to hit Wrong answer on Combined, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
63/63
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #104 | Gemini 3.1 Flash Lite Preview low | 1 | 3.0 | $0.646 | 0/2 | 160.6s | |
| #52 | Kimi K2.7 Code medium | Moonshot AI | 1 | 7.3 | $0.751 | 1/2 | 66.0s |
| #20 | Grok 4.5 low | X AI | 1 | 6.5 | $0.935 | 1/2 | 12.8s |