Agentic Ranking
See which AI models perform best on Agentic, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↓.
380/380
Filter models
No models match the current search and filters.
| Rank | Model | Company | Agentic Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #255 | GPT-5.4 Mini none | OpenAI | 5.0 | 5.7 | $0.170 | 0/1 | 42.0s |
| #256 | Seed 2.1 Turbo none | Bytedance Seed | 5.0 | 5.7 | $0.182 | 0/1 | 71.4s |
| #257 | Dots 3 Note Preview medium | Dots Studio | 4.7 | 5.7 | $0.000 | 0/1 | 140.5s |
| #258 | Gemma 4 31B none | 3.9 | 5.7 | $0.030 | 0/1 | 44.7s | |
| #259 | Space Bunny Alpha low | Stealth | 4.7 | 5.7 | $0.000 | 0/1 | 89.0s |
| #260 | GPT-5.4 none | OpenAI | 5.0 | 5.6 | $0.690 | 0/1 | 43.7s |
| #261 | Solar Pro 4 none | Upstage | 5.0 | 5.6 | $0.033 | 0/1 | 69.2s |
| #262 | GLM 5 none | Z.ai | 5.0 | 5.6 | $0.126 | 0/1 | 74.8s |
| #263 | Qwen3.5-35B-A3B none | Qwen | 3.9 | 5.6 | $0.213 | 0/1 | 141.2s |
| #264 | Nemotron 3 Super medium | NVIDIA | 4.7 | 5.5 | $0.081 | 0/1 | 138.3s |
| #265 | Kimi K2.6 none | Moonshot AI | 4.7 | 5.5 | $0.376 | 0/1 | 175.2s |
| #266 | Gemini 2.5 Flash none | 3.0 | 5.5 | $0.017 | 0/1 | 1.81s | |
| #267 | GPT-6 Luna none | OpenAI | 5.0 | 5.5 | $0.026 | 0/1 | 51.1s |
| #268 | Gemma 4 26B A4B none | 5.0 | 5.5 | $0.026 | 0/1 | 75.8s | |
| #269 | GLM 5V Turbo medium | Z.ai | 0.0 | 5.5 | $0.457 | 0/0 | 0ms |