AI BENCHY
Advertise here

Kushindwa kwa AI BENCHY

Kushindwa kwa Jibu lisilo sahihi

Ona ni modeli gani za AI hukutana na Jibu lisilo sahihi mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Muda wa majibu (wastani) ↑.

Modeli zilizoonyeshwa

15

Jumla ya kushindwa

1204

Modeli iliyoathirika zaidi

Mistral Small 4 15
Nafasi Modeli Kampuni Idadi ya Jibu lisilo sahihi Alama Majaribio sahihi Muda wa majibu (wastani)
#87 Gemini 3.1 Flash Lite minimal Google 8 6.4 10/21 1.33s
#125 GPT-5.4 none OpenAI 13 5.5 7/21 1.42s
#148 GPT-5.4 Nano none OpenAI 15 4.7 4/21 1.48s
#32 Gemini 3.5 Flash minimal Google 5 7.7 14/21 1.57s
#128 Qwen3.6 Flash none Qwen 12 5.4 7/21 1.60s
#157 Grok 4.1 Fast none X AI 13 4.4 3/19 1.62s
#48 Gemini 3 Flash Preview none Google 8 7.4 13/21 1.65s
#115 Qwen3.5-27B none Qwen 12 5.7 7/21 1.68s
#147 GPT-4o-mini none OpenAI 15 4.8 5/21 1.77s
#123 MiMo-V2.5-Pro none Xiaomi 11 5.5 6/21 1.78s
#61 Gemini 3.1 Flash Lite low Google 9 7.2 12/21 1.89s
#91 GPT-5.5 none OpenAI 11 6.4 10/21 1.89s
#154 Qwen3.5-9B none Qwen 14 4.6 4/21 1.89s
#143 MiMo-V2.5 none Xiaomi 14 4.9 5/21 2.20s
#81 Mercury 2 medium Inception 8 6.6 10/21 2.24s

Modeli bora kwa Idadi ya Jibu lisilo sahihi

Idadi ya Jibu lisilo sahihi dhidi ya Alama

Modeli bora kwa Muda wa majibu (wastani)