Classement Tâches agentiques
Voyez quels modèles d'IA réussissent le mieux sur Tâches agentiques, lesquels restent fiables et où les écarts sont les plus marqués. Trier par: Coût total ↑.
380/380
Filtrer les modèles
Aucun modèle ne correspond à la recherche et aux filtres actuels.
| Rang | Modèle | Entreprise | Score Tâches agentiques | Score | Coût total | Tests corrects | Temps de réponse (moy.) |
|---|---|---|---|---|---|---|---|
| #235 | Granite 4.2 8B medium | IBM Granite | 10.0 | 6.1 | $0.029 | 1/1 | 149.5s |
| #243 | gpt-oss-120b medium | OpenAI | 5.0 | 5.9 | $0.030 | 0/1 | 203.9s |
| #258 | Gemma 4 31B none | 3.9 | 5.7 | $0.030 | 0/1 | 44.7s | |
| #115 | GPT-6 Luna low | OpenAI | 10.0 | 7.6 | $0.033 | 1/1 | 51.0s |
| #350 | Laguna M.1 medium | Poolside | 0.0 | 3.9 | $0.033 | 0/0 | 0ms |
| #261 | Solar Pro 4 none | Upstage | 5.0 | 5.6 | $0.033 | 0/1 | 69.2s |
| #338 | Mercury 2 none | Inception | 3.0 | 4.4 | $0.033 | 0/1 | 9.00s |
| #207 | Mercury 2.5 medium | Inception | 4.7 | 6.4 | $0.034 | 0/1 | 107.9s |
| #280 | Granite 4.2 8B low | IBM Granite | 6.1 | 5.3 | $0.036 | 0/1 | 142.7s |
| #171 | Mercury 2.5 Preview medium | Inception | 3.9 | 6.8 | $0.037 | 0/1 | 74.6s |
| #193 | Mercury 2.5 Preview high | Inception | 6.1 | 6.5 | $0.039 | 0/1 | 86.8s |
| #309 | Qwen3.5-9B none | Qwen | 3.9 | 4.9 | $0.041 | 0/1 | 433.0s |
| #330 | Hy3 preview low | Tencent | 0.0 | 4.5 | $0.042 | 0/0 | 0ms |
| #290 | MiMo-V2-Flash medium | Xiaomi | 0.0 | 5.1 | $0.043 | 0/0 | 0ms |
| #194 | Mercury 2.5 high | Inception | 3.9 | 6.5 | $0.044 | 0/1 | 80.3s |