Categorie AI BENCHY
Clasament Respectarea instrucțiunilor
Vezi ce modele AI se descurcă cel mai bine la Respectarea instrucțiunilor, care rămân fiabile și unde apar cele mai mari diferențe.
Modele afișate
15
Media pentru Scor Respectarea instrucțiunilor
8.5
Cel mai bun model
Gemini 3 Flash Preview 10.0| Rang | Model | Companie | Scor Respectarea instrucțiunilor | Scor | Teste corecte | Timp de răspuns (mediu) |
|---|---|---|---|---|---|---|
| #145 | Laguna M.1 none | Poolside | 6.3 | 4.8 | 1/2 | 683ms |
| #156 | Hy3 preview none | Tencent | 6.3 | 4.4 | 1/2 | 13.0s |
| #88 | Qwen3.7 Plus none | Qwen | 6.3 | 6.4 | 1/2 | 929ms |
| #102 | Gemma 4 26B A4B none | 6.3 | 6.0 | 1/2 | 690ms | |
| #106 | Grok 4.20 Beta none | X AI | 6.3 | 5.8 | 1/2 | 649ms |
| #108 | Qwen3.5-Flash none | Qwen | 6.3 | 5.8 | 1/2 | 8.81s |
| #115 | Qwen3.5-27B none | Qwen | 6.3 | 5.7 | 1/2 | 1.03s |
| #117 | Qwen3.5-35B-A3B none | Qwen | 6.3 | 5.6 | 1/2 | 809ms |
| #127 | Grok 4.20 none | X AI | 6.3 | 5.4 | 1/2 | 445ms |
| #128 | Qwen3.6 Flash none | Qwen | 6.3 | 5.4 | 1/2 | 1.10s |
| #131 | Qwen3.5-122B-A10B none | Qwen | 6.3 | 5.3 | 1/2 | 513ms |
| #140 | Qwen3 Coder Next none | Qwen | 6.3 | 4.9 | 1/2 | 7.78s |
| #141 | Nemotron 3 Super none | NVIDIA | 6.3 | 4.9 | 1/2 | 804ms |
| #144 | GPT-5.4 Mini none | OpenAI | 6.3 | 4.9 | 1/2 | 728ms |
| #147 | GPT-4o-mini none | OpenAI | 6.3 | 4.8 | 1/2 | 1.11s |