General Intelligence: Wrong answer
General Intelligence
Wrong answer
See which AI models are most likely to hit Wrong answer on General Intelligence, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
95/95
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #55 | Grok 4.5 medium | X AI | 1 | 6.5 | $2.042 | 0/1 | 12.8s |
| #208 | Nemotron 3 Ultra none | NVIDIA | 1 | 5.0 | $0.093 | 0/1 | 13.5s |
| #286 | Ring-2.6-1T none | Inclusionai | 1 | 4.3 | $0.026 | 0/1 | 15.6s |
| #117 | LongCat 2.0 medium | Meituan | 1 | 4.8 | $0.499 | 0/1 | 16.4s |
| #299 | Granite 4.2 8B high | IBM Granite | 1 | 5.5 | $0.135 | 0/1 | 16.4s |
| #170 | LongCat 2.0 high | Meituan | 1 | 5.1 | $0.492 | 0/1 | 17.0s |
| #57 | Grok 4.6 medium | X AI | 1 | 4.1 | $1.713 | 0/1 | 17.3s |
| #306 | GLM 4.7 Flash medium | Z.ai | 1 | 3.6 | $0.168 | 0/1 | 18.1s |
| #111 | Grok Build 0.1 medium | X AI | 1 | 4.4 | $1.130 | 0/1 | 18.4s |
| #186 | Dots 3 Note Preview high | Dots Studio | 1 | 3.4 | $0.000 | 0/1 | 19.6s |
| #167 | LongCat 2.0 low | Meituan | 1 | 3.4 | $0.412 | 0/1 | 22.5s |
| #69 | Qwen3.8 2.4T A95B low | Qwen | 1 | 6.1 | $2.672 | 0/1 | 23.0s |
| #248 | DeepSeek V4 Flash 0423 none | DeepSeek | 1 | 4.2 | $0.040 | 0/1 | 23.7s |
| #226 | North Mini Code medium | Cohere | 1 | 5.1 | $0.000 | 0/1 | 25.1s |
| #132 | Qwen3.5 Plus 2026-04-20 medium | Qwen | 1 | 4.9 | $0.323 | 0/1 | 25.3s |