AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Kushindwa kwa AI BENCHY

Kushindwa kwa Jibu lisilo sahihi

Ona ni modeli gani za AI hukutana na Jibu lisilo sahihi mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Muda wa majibu (wastani) ↑.

Modeli zilizoonyeshwa

15

Jumla ya kushindwa

1204

Modeli iliyoathirika zaidi

Mistral Small 4 15
Nafasi Modeli Kampuni Idadi ya Jibu lisilo sahihi Alama Majaribio sahihi Muda wa majibu (wastani)
#3 Gemini 3.5 Flash low Google 2 9.4 19/21 3.27s
#74 Qwen3.6 Max Preview none Qwen 10 6.9 11/21 3.30s
#117 Qwen3.5-35B-A3B none Qwen 12 5.6 7/21 3.37s
#131 Qwen3.5-122B-A10B none Qwen 13 5.3 6/21 3.41s
#68 Claude Opus 4.8 none Anthropic 4 7.0 12/21 3.47s
#108 Qwen3.5-Flash none Qwen 13 5.8 8/21 3.58s
#118 Qwen3.6 27B none Qwen 11 5.6 7/21 3.72s
#153 Qwen3.6 35B A3B none Qwen 13 4.6 4/21 3.73s
#40 Gemini 3.1 Flash Lite Preview medium Google 7 7.5 13/21 3.96s
#98 GLM 5 none Z.ai 12 6.1 9/21 4.03s
#85 Gemma 4 31B none Google 8 6.5 10/21 4.05s
#112 GLM 5.1 none Z.ai 13 5.7 7/21 4.10s
#114 Qwen3.5 Plus 2026-04-20 none Qwen 12 5.7 7/21 4.39s
#116 Hunter Alpha none OpenRouter 9 5.7 6/18 4.70s
#11 Claude Opus 4.7 medium Anthropic 3 8.7 17/21 4.73s

Modeli bora kwa Idadi ya Jibu lisilo sahihi

Idadi ya Jibu lisilo sahihi dhidi ya Alama

Modeli bora kwa Muda wa majibu (wastani)