Agentic Ranking
See which AI models perform best on Agentic, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↓.
380/380
Filter models
No models match the current search and filters.
| Rank | Model | Company | Agentic Score | Score | Total Cost | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|---|
| #315 | Nemotron 3.5 Lightning medium | NVIDIA | 2.8 | 4.7 | $0.129 | 0/1 | 341.8s |
| #316 | North Mini Code none | Cohere | 3.0 | 4.7 | $0.000 | 0/1 | 18.9s |
| #317 | Laguna S 2.1 none | Poolside | 5.4 | 4.7 | $0.047 | 0/1 | 149.3s |
| #318 | Granite 4.2 8B high | IBM Granite | 5.0 | 4.7 | $0.161 | 0/1 | 223.5s |
| #319 | Qwen3 Coder Next medium | Qwen | 5.0 | 4.7 | $0.069 | 0/1 | 151.4s |
| #320 | DeepSeek V3.2 none | DeepSeek | 3.0 | 4.6 | $0.121 | 0/1 | 122.5s |
| #321 | GPT-5.4 Nano none | OpenAI | 3.9 | 4.6 | $0.068 | 0/1 | 45.4s |
| #322 | Gemini 3.1 Flash Lite high | 0.0 | 4.6 | $2.044 | 0/0 | 0ms | |
| #323 | GLM 5V Turbo none | Z.ai | 0.0 | 4.6 | $0.052 | 0/0 | 0ms |
| #324 | Laguna S 2.1 low | Poolside | 2.8 | 4.6 | $0.109 | 0/1 | 152.1s |
| #325 | Nemotron 3.5 Lightning low | NVIDIA | 3.5 | 4.6 | $0.117 | 0/1 | 597.6s |
| #326 | Owl Alpha medium | Openrouter | 0.0 | 4.6 | $0.000 | 0/0 | 0ms |
| #327 | Mimo V2 PRO none | Xiaomi | 0.0 | 4.6 | $0.045 | 0/0 | 0ms |
| #328 | Owl Alpha none | Openrouter | 0.0 | 4.6 | $0.000 | 0/0 | 0ms |
| #329 | Ling 2.6 Flash none | Inclusionai | 3.0 | 4.5 | $0.002 | 0/1 | 32ms |