AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Kushindwa kwa AI BENCHY

Kushindwa kwa Jibu lisilo sahihi

Ona ni modeli gani za AI hukutana na Jibu lisilo sahihi mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Muda wa majibu (wastani) ↓.

Modeli zilizoonyeshwa

15

Jumla ya kushindwa

1204

Modeli iliyoathirika zaidi

Kimi K2.5 5
Nafasi Modeli Kampuni Idadi ya Jibu lisilo sahihi Alama Majaribio sahihi Muda wa majibu (wastani)
#114 Qwen3.5 Plus 2026-04-20 none Qwen 12 5.7 7/21 4.39s
#112 GLM 5.1 none Z.ai 13 5.7 7/21 4.10s
#85 Gemma 4 31B none Google 8 6.5 10/21 4.05s
#98 GLM 5 none Z.ai 12 6.1 9/21 4.03s
#40 Gemini 3.1 Flash Lite Preview medium Google 7 7.5 13/21 3.96s
#153 Qwen3.6 35B A3B none Qwen 13 4.6 4/21 3.73s
#118 Qwen3.6 27B none Qwen 11 5.6 7/21 3.72s
#108 Qwen3.5-Flash none Qwen 13 5.8 8/21 3.58s
#68 Claude Opus 4.8 none Anthropic 4 7.0 12/21 3.47s
#131 Qwen3.5-122B-A10B none Qwen 13 5.3 6/21 3.41s
#117 Qwen3.5-35B-A3B none Qwen 12 5.6 7/21 3.37s
#74 Qwen3.6 Max Preview none Qwen 10 6.9 11/21 3.30s
#3 Gemini 3.5 Flash low Google 2 9.4 19/21 3.27s
#44 Gemini 3.1 Flash Lite medium Google 7 7.5 13/21 3.23s
#8 Claude Opus 4.7 none Anthropic 3 8.9 16/19 3.02s

Modeli bora kwa Idadi ya Jibu lisilo sahihi

Idadi ya Jibu lisilo sahihi dhidi ya Alama

Modeli bora kwa Muda wa majibu (wastani)