Agentic: Wrong answer
Agentic
Wrong answer
See which AI models are most likely to hit Wrong answer on Agentic, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
108/108
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #342 | gpt-oss-120b none | OpenAI | 1 | 6.1 | $0.016 | 0/1 | 140.9s |
| #263 | Qwen3.5-35B-A3B none | Qwen | 1 | 3.9 | $0.213 | 0/1 | 141.2s |
| #280 | Granite 4.2 8B low | IBM Granite | 1 | 6.1 | $0.036 | 0/1 | 142.7s |
| #187 | LongCat 2.0 low | Meituan | 1 | 6.1 | $0.495 | 0/1 | 145.6s |
| #319 | Qwen3 Coder Next medium | Qwen | 1 | 5.0 | $0.069 | 0/1 | 151.4s |
| #221 | Space Bunny Alpha medium | Stealth | 1 | 6.1 | $0.000 | 0/1 | 153.1s |
| #127 | Grok 4.7 low | X AI | 1 | 5.0 | $2.263 | 0/1 | 154.9s |
| #368 | Command A+ high | Cohere | 1 | 3.9 | $0.079 | 0/1 | 156.0s |
| #305 | Laguna S 2.1 high | Poolside | 1 | 2.8 | $0.147 | 0/1 | 158.2s |
| #62 | Grok 4.6 high | X AI | 1 | 5.0 | $2.629 | 0/1 | 160.9s |
| #211 | Dots 3 Note Preview low | Dots Studio | 1 | 6.1 | $0.000 | 0/1 | 161.2s |
| #296 | Laguna S 2.1 medium | Poolside | 1 | 3.4 | $0.080 | 0/1 | 165.0s |
| #276 | Step 3.5 Flash medium | Stepfun | 1 | 2.8 | $0.105 | 0/1 | 172.2s |
| #144 | LongCat 2.0 medium | Meituan | 1 | 6.1 | $0.593 | 0/1 | 174.1s |
| #265 | Kimi K2.6 none | Moonshot AI | 1 | 4.7 | $0.376 | 0/1 | 175.2s |