Kushindwa kwa AI BENCHY
Kushindwa kwa Muundo wa ziada
Ona ni modeli gani za AI hukutana na Muundo wa ziada mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Idadi ya kushindwa ↑.
32/32
Chuja miundo
Hakuna miundo inayolingana na utafutaji na vichujio vya sasa.
| Nafasi | Modeli | Kampuni | Idadi ya Muundo wa ziada | Alama | Jumla ya gharama | Majaribio sahihi | Muda wa majibu (wastani) |
|---|---|---|---|---|---|---|---|
| #29 | Qwen3.5-27B medium | Qwen | 1 | 7.9 | $0.536 | 13/21 | 68.4s |
| #37 | Grok 4.3 medium | X AI | 1 | 7.7 | $0.614 | 13/21 | 47.5s |
| #40 | MiniMax M3 medium | Minimax | 1 | 7.6 | $0.131 | 11/21 | 68.2s |
| #41 | DeepSeek V4 Pro high | DeepSeek | 1 | 7.6 | $0.157 | 9/21 | 77.2s |
| #53 | Grok 4.20 medium | X AI | 1 | 7.3 | $0.609 | 12/21 | 27.7s |
| #58 | DeepSeek V4 Pro none | DeepSeek | 1 | 7.2 | $0.034 | 10/21 | 6.41s |
| #62 | MiMo-V2-Flash medium | Xiaomi | 1 | 7.1 | $0.043 | 12/21 | 20.1s |
| #64 | GLM 5.1 medium | Z.ai | 1 | 7.1 | $0.292 | 12/21 | 33.7s |
| #73 | Mimo V2 Omni medium | Xiaomi | 1 | 6.8 | $0.683 | 10/21 | 41.2s |
| #77 | Mimo V2 PRO medium | Xiaomi | 1 | 6.7 | $0.333 | 12/21 | 22.2s |
| #110 | Owl Alpha none | Openrouter | 1 | 5.8 | $0.000 | 7/21 | 9.88s |
| #114 | Mimo V2 Omni none | Xiaomi | 1 | 5.7 | $0.021 | 8/21 | 2.44s |
| #130 | Qwen3 Coder Next none | Qwen | 1 | 5.1 | $0.009 | 5/21 | 8.62s |
| #132 | Hunter Alpha medium | OpenRouter | 1 | 5.1 | $0.000 | 8/18 | 10.3s |
| #134 | MiMo-V2.5 none | Xiaomi | 1 | 5.1 | $0.007 | 5/21 | 2.20s |