Agentic: Wrong answer
Agentic
Wrong answer
See which AI models are most likely to hit Wrong answer on Agentic, so you can spot weak points faster. Sort by: Total Cost ↑.
Failure Reasons
108/108
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #204 | Solar Pro 4 medium | Upstage | 1 | 4.7 | $0.174 | 0/1 | 272.3s |
| #256 | Seed 2.1 Turbo none | Bytedance Seed | 1 | 5.0 | $0.182 | 0/1 | 71.4s |
| #215 | Gemini 3.5 Flash Lite low | 1 | 5.0 | $0.183 | 0/1 | 53.9s | |
| #341 | GLM 4.7 Flash medium | Z.ai | 1 | 3.9 | $0.193 | 0/1 | 399.2s |
| #247 | Nemotron 3 Ultra none | NVIDIA | 1 | 5.0 | $0.193 | 0/1 | 70.2s |
| #151 | MiMo-V2.6-Flash low | Xiaomi | 1 | 4.7 | $0.200 | 0/1 | 107.9s |
| #263 | Qwen3.5-35B-A3B none | Qwen | 1 | 3.9 | $0.213 | 0/1 | 141.2s |
| #281 | Kimi K2.5 none | Moonshot AI | 1 | 5.0 | $0.236 | 0/1 | 321.2s |
| #160 | Gemini 3 Flash Preview low | 1 | 5.0 | $0.254 | 0/1 | 52.6s | |
| #270 | Seed-2.0-Code none | Bytedance Seed | 1 | 6.1 | $0.257 | 0/1 | 80.0s |
| #124 | DeepSeek V4 Flash 0731 low | DeepSeek | 1 | 4.7 | $0.307 | 0/1 | 321.2s |
| #145 | GLM 5 medium | Z.ai | 1 | 4.7 | $0.360 | 0/1 | 110.7s |
| #180 | MiniMax M3 medium | Minimax | 1 | 3.0 | $0.372 | 0/1 | 116.5s |
| #265 | Kimi K2.6 none | Moonshot AI | 1 | 4.7 | $0.376 | 0/1 | 175.2s |
| #137 | Qwen3.6 Plus medium | Qwen | 1 | 4.7 | $0.391 | 0/1 | 307.1s |