Kushindwa kwa AI BENCHY
Kushindwa kwa Hakufuata maelekezo
Ona ni modeli gani za AI hukutana na Hakufuata maelekezo mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Majaribio sahihi ↓.
| Nafasi | Modeli | Kampuni | Idadi ya Hakufuata maelekezo | Alama | Majaribio sahihi | Muda wa majibu (wastani) |
|---|---|---|---|---|---|---|
| #48 | Gemma 4 31B none | 1 | 6.9 | 10/18 | 4.02s | |
| #44 | GPT-5.4 Mini medium | OpenAI | 5 | 7.3 | 9/18 | 15.2s |
| #45 | GPT-5 Mini medium | OpenAI | 4 | 7.0 | 9/18 | 24.0s |
| #46 | Kimi K2.5 medium | Moonshot AI | 2 | 7.0 | 9/18 | 72.4s |
| #47 | Grok 4.20 medium | X AI | 4 | 7.0 | 9/18 | 10.3s |
| #51 | Nemotron 3 Super medium | NVIDIA | 4 | 6.7 | 9/18 | 19.1s |
| #52 | Grok 4.1 Fast medium | X AI | 4 | 6.7 | 9/18 | 23.9s |
| #50 | Hunter Alpha medium | OpenRouter | 2 | 6.7 | 8/18 | 10.3s |
| #54 | Mercury 2 medium | Inception | 4 | 6.5 | 8/18 | 2.21s |
| #55 | MiMo-V2-Omni none | Xiaomi | 2 | 6.5 | 8/18 | 1.99s |
| #58 | GLM 5V Turbo none | Z.ai | 2 | 6.2 | 8/18 | 3.10s |
| #59 | Qwen3.5-Flash none | Qwen | 1 | 6.2 | 8/18 | 3.25s |
| #56 | Grok 4.20 Multi Agent Beta medium | X AI | 4 | 6.4 | 7/18 | 9.80s |
| #57 | GPT-5 Nano medium | OpenAI | 3 | 6.3 | 7/18 | 44.1s |
| #60 | Gemma 4 26B A4B none | 3 | 6.2 | 7/18 | 6.59s |