General Intelligence: Wrong answer
General Intelligence
Wrong answer
See which AI models are most likely to hit Wrong answer on General Intelligence, so you can spot weak points faster. Sort by: Response Time (avg) ↑.
Failure Reasons
62/62
Filter models
No models match the current search and filters.
| Rank | Model | Company | Wrong answer Count | Category Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #89 | Qwen3.6 Flash medium | Qwen | 1 | 4.8 | $0.738 | 0/1 | 9.88s |
| #53 | GLM 5 Turbo medium | Z.ai | 1 | 6.1 | $0.323 | 0/1 | 10.1s |
| #212 | gpt-oss-120b none | OpenAI | 1 | 4.8 | $0.010 | 0/1 | 10.8s |
| #188 | KAT-Coder-Air V2.5 none | Kwaipilot | 1 | 5.0 | $0.067 | 0/1 | 12.0s |
| #208 | Grok Build 0.1 none | X AI | 1 | 4.3 | $0.547 | 0/1 | 12.5s |
| #25 | Grok 4.5 medium | X AI | 1 | 6.5 | $1.928 | 0/1 | 12.8s |
| #135 | Nemotron 3 Ultra none | NVIDIA | 1 | 5.0 | $0.095 | 0/1 | 13.5s |
| #185 | Ring-2.6-1T none | Inclusionai | 1 | 4.3 | $0.026 | 0/1 | 15.6s |
| #64 | LongCat 2.0 medium | Meituan | 1 | 4.8 | $0.478 | 0/1 | 16.4s |
| #102 | LongCat 2.0 high | Meituan | 1 | 5.1 | $0.469 | 0/1 | 17.0s |
| #200 | GLM 4.7 Flash medium | Z.ai | 1 | 3.6 | $0.166 | 0/1 | 18.1s |
| #52 | Grok Build 0.1 medium | X AI | 1 | 4.4 | $1.097 | 0/1 | 18.4s |
| #96 | LongCat 2.0 low | Meituan | 1 | 3.4 | $0.391 | 0/1 | 22.5s |
| #156 | DeepSeek V4 Flash none | DeepSeek | 1 | 4.2 | $0.042 | 0/1 | 23.7s |
| #143 | North Mini Code medium | Cohere | 1 | 5.1 | $0.000 | 0/1 | 25.1s |