Kategori AI BENCHY
Peringkat Kepatuhan instruksi
Lihat model AI mana yang paling baik di Kepatuhan instruksi, mana yang tetap andal, dan di mana kesenjangan terbesar muncul. Urutkan berdasarkan: Metrik ↑.
| Peringkat | Model | Perusahaan | Skor Kepatuhan instruksi | Skor | Tes benar | Waktu respons (rata-rata) |
|---|---|---|---|---|---|---|
| #124 | Kimi K2.6 none | Moonshot AI | 6.5 | 5.5 | 1/2 | 1.64s |
| #125 | GPT-5.4 none | OpenAI | 6.5 | 5.5 | 1/2 | 1.07s |
| #134 | GLM 5 Turbo none | Z.ai | 6.5 | 5.2 | 1/2 | 2.13s |
| #135 | Kimi K2.5 none | Moonshot AI | 6.5 | 5.2 | 1/2 | 2.67s |
| #139 | DeepSeek V4 Flash none | DeepSeek | 6.5 | 5.0 | 1/2 | 17.5s |
| #142 | Mistral Small 4 none | Mistral | 6.5 | 4.9 | 1/2 | 380ms |
| #143 | MiMo-V2.5 none | Xiaomi | 6.5 | 4.9 | 1/2 | 751ms |
| #146 | Laguna Xs.2 none | Poolside | 6.5 | 4.8 | 1/2 | 439ms |
| #152 | MiMo-V2-Flash none | Xiaomi | 6.5 | 4.6 | 1/2 | 857ms |
| #154 | Qwen3.5-9B none | Qwen | 6.5 | 4.6 | 1/2 | 514ms |
| #155 | Mercury 2 none | Inception | 6.5 | 4.5 | 1/2 | 551ms |
| #161 | Qwen3.5-9B medium | Qwen | 6.5 | 4.2 | 1/2 | 5.75s |
| #105 | Nemotron 3 Super medium | NVIDIA | 7.3 | 5.8 | 1/2 | 6.97s |
| #149 | Nemotron 3 Nano Omni 30b A3b Reasoning medium | NVIDIA | 7.3 | 4.6 | 1/2 | 1.37s |
| #53 | Gemini 3.1 Flash Lite high | 7.3 | 7.3 | 1/2 | 23.3s |