Kushindwa kwa kategoria za AI BENCHY
Mbinu za kupinga AI: Hakufuata maelekezo
Mbinu za kupinga AI
Hakufuata maelekezo
Ona ni modeli gani za AI zina uwezekano mkubwa wa kupata Hakufuata maelekezo katika Mbinu za kupinga AI, ili uone udhaifu haraka. Panga kwa: Majaribio sahihi ↑.
Sababu za kushindwa
29/29
Chuja miundo
Hakuna miundo inayolingana na utafutaji na vichujio vya sasa.
| Nafasi | Modeli | Kampuni | Idadi ya Hakufuata maelekezo | Alama ya kategoria | Jumla ya gharama | Majaribio sahihi | Muda wa majibu (wastani) |
|---|---|---|---|---|---|---|---|
| #58 | DeepSeek V4 Pro none | DeepSeek | 1 | 3.2 | $0.034 | 0/4 | 4.02s |
| #110 | Owl Alpha none | Openrouter | 1 | 3.4 | $0.000 | 0/4 | 2.78s |
| #114 | Mimo V2 Omni none | Xiaomi | 1 | 3.6 | $0.021 | 0/4 | 1.63s |
| #119 | MiMo-V2.5-Pro none | Xiaomi | 1 | 3.3 | $0.017 | 0/4 | 2.67s |
| #130 | Qwen3 Coder Next none | Qwen | 1 | 3.6 | $0.009 | 0/4 | 3.31s |
| #148 | Qwen3 Coder Next medium | Qwen | 1 | 3.5 | $0.008 | 0/4 | 8.64s |
| #161 | Grok 4.1 Fast none | X AI | 1 | 3.2 | $0.008 | 0/4 | 1.07s |
| #162 | Laguna Xs.2 none | Poolside | 1 | 3.0 | $0.000 | 0/4 | 534ms |
| #40 | MiniMax M3 medium | Minimax | 1 | 5.5 | $0.131 | 1/4 | 14.9s |
| #157 | GLM 4.7 Flash medium | Z.ai | 1 | 4.7 | $0.054 | 1/4 | 15.0s |
| #158 | Hy3 preview none | Tencent | 2 | 4.8 | $0.003 | 1/4 | 11.1s |
| #163 | Granite 4.1 8B none | IBM Granite | 1 | 4.9 | $0.003 | 1/4 | 844ms |
| #16 | GPT-5 Mini medium | OpenAI | 1 | 7.1 | $0.159 | 2/4 | 13.9s |
| #22 | GPT-5.2 medium | OpenAI | 1 | 6.5 | $0.548 | 2/4 | 7.81s |
| #35 | Kimi K2.6 medium | Moonshot AI | 1 | 7.0 | $0.889 | 2/4 | 11.6s |