AI BENCHY
Advertise here

Kushindwa kwa AI BENCHY

Kushindwa kwa Mwito wa zana si sahihi

Ona ni modeli gani za AI hukutana na Mwito wa zana si sahihi mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Alama ↑.

Modeli zilizoonyeshwa

15

Jumla ya kushindwa

26

Modeli iliyoathirika zaidi

Granite 4.1 8B 1
Nafasi Modeli Kampuni Idadi ya Mwito wa zana si sahihi Alama Majaribio sahihi Muda wa majibu (wastani)
#163 Granite 4.1 8B none IBM Granite 1 4.0 2/21 728ms
#159 Ling-2.6-1T none Inclusionai 1 4.3 3/21 7.72s
#158 GLM 4.7 Flash medium Z.ai 1 4.4 4/21 35.1s
#154 Qwen3.5-9B none Qwen 1 4.6 4/21 1.89s
#146 Laguna Xs.2 none Poolside 1 4.8 5/19 806ms
#145 Laguna M.1 none Poolside 1 4.8 4/19 2.89s
#139 DeepSeek V4 Flash none DeepSeek 1 5.0 5/21 26.8s
#138 Ling-2.6-flash none Inclusionai 2 5.0 6/21 9.34s
#137 Elephant Alpha none Openrouter 1 5.1 5/21 1.22s
#136 Elephant Alpha medium Openrouter 1 5.1 6/21 1.27s
#133 DeepSeek V3.2 none DeepSeek 1 5.2 6/21 13.8s
#130 MiniMax M2.7 medium Minimax 1 5.3 5/21 38.2s
#129 MiniMax M2.5 medium Minimax 1 5.3 5/21 65.4s
#128 Qwen3.6 Flash none Qwen 1 5.4 7/21 1.60s
#127 Grok 4.20 none X AI 1 5.4 6/18 1.11s

Modeli bora kwa Idadi ya Mwito wa zana si sahihi

Idadi ya Mwito wa zana si sahihi dhidi ya Alama

Modeli bora kwa Muda wa majibu (wastani)