Kushindwa kwa AI BENCHY
Kushindwa kwa Hakufuata maelekezo
Ona ni modeli gani za AI hukutana na Hakufuata maelekezo mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Muda wa majibu (wastani) ↓.
| Nafasi | Modeli | Kampuni | Idadi ya Hakufuata maelekezo | Alama | Majaribio sahihi | Muda wa majibu (wastani) |
|---|---|---|---|---|---|---|
| #14 | Gemma 4 31B medium | 1 | 8.3 | 13/18 | 24.9s | |
| #45 | GPT-5 Mini medium | OpenAI | 4 | 7.0 | 9/18 | 24.0s |
| #52 | Grok 4.1 Fast medium | X AI | 4 | 6.7 | 9/18 | 23.9s |
| #41 | MiMo-V2-Flash medium | Xiaomi | 1 | 7.5 | 11/18 | 23.4s |
| #13 | GLM 5 medium | Z.ai | 1 | 8.4 | 13/18 | 23.3s |
| #51 | Nemotron 3 Super medium | NVIDIA | 4 | 6.7 | 9/18 | 19.1s |
| #16 | GPT-5.4 medium | OpenAI | 2 | 8.2 | 13/18 | 18.6s |
| #18 | GLM 5 Turbo medium | Z.ai | 2 | 8.1 | 12/18 | 17.7s |
| #35 | MiMo-V2-Omni medium | Xiaomi | 2 | 7.7 | 11/18 | 16.8s |
| #68 | gpt-oss-120b medium | OpenAI | 4 | 5.8 | 7/18 | 16.1s |
| #7 | GPT-5.3-Codex medium | OpenAI | 2 | 8.6 | 13/18 | 15.4s |
| #20 | Qwen3.6 Plus medium | Qwen | 1 | 8.1 | 13/18 | 15.3s |
| #44 | GPT-5.4 Mini medium | OpenAI | 5 | 7.3 | 9/18 | 15.2s |
| #31 | GLM 5V Turbo medium | Z.ai | 2 | 7.8 | 11/18 | 15.0s |
| #40 | GPT-5.2 medium | OpenAI | 3 | 7.5 | 11/18 | 14.0s |