General Intelligence Ranking
See which AI models perform best on General Intelligence, which ones stay reliable, and where the biggest gaps appear. Sort by: Metric ↑.
319/319
Filter models
No models match the current search and filters.
| Rank | Model | Company | General Intelligence Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #256 | Inkling none | Thinkingmachines | 5.0 | 5.2 | $0.147 | 0/1 | 859ms |
| #268 | Mercury 2.5 Preview low | Inception | 5.0 | 5.0 | $0.011 | 0/1 | 1.05s |
| #273 | Mercury 2.5 Preview none | Inception | 5.0 | 4.9 | $0.018 | 0/1 | 482ms |
| #282 | KAT-Coder-Air V2.5 none | Kwaipilot | 5.0 | 4.8 | $0.070 | 0/1 | 12.0s |
| #293 | Laguna S 2.1 none | Poolside | 5.0 | 4.5 | $0.022 | 0/1 | 815ms |
| #294 | Grok 4.20 Beta none | X AI | 5.0 | 4.4 | $0.087 | 0/1 | 541ms |
| #299 | Granite 4.2 8B none | IBM Granite | 5.0 | 4.3 | $0.026 | 0/1 | 770ms |
| #227 | Gemini 3.1 Flash Lite high | 5.0 | 5.6 | $2.044 | 0/1 | 45.7s | |
| #73 | GPT-5.6 Terra high | OpenAI | 5.1 | 8.0 | $0.836 | 0/1 | 3.03s |
| #79 | Qwen3.6 Plus medium | Qwen | 5.1 | 7.8 | $0.418 | 0/1 | 27.1s |
| #111 | GPT-5.6 Luna medium | OpenAI | 5.1 | 7.4 | $0.072 | 0/1 | 4.34s |
| #125 | KAT-Coder-Pro V2.5 high | Kwaipilot | 5.1 | 7.2 | $0.506 | 0/1 | 3.27s |
| #163 | LongCat 2.0 high | Meituan | 5.1 | 6.7 | $0.492 | 0/1 | 17.0s |
| #171 | Mercury 2.5 Preview high | Inception | 5.1 | 6.6 | $0.030 | 0/1 | 1.88s |
| #219 | North Mini Code medium | Cohere | 5.1 | 5.7 | $0.000 | 0/1 | 25.1s |