AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Kushindwa kwa AI BENCHY

Kushindwa kwa Jibu lisilo sahihi

Ona ni modeli gani za AI hukutana na Jibu lisilo sahihi mara nyingi zaidi ili utambue hatari za utegemevu kabla ya kuchagua. Panga kwa: Muda wa majibu (wastani) ↓.

Modeli zilizoonyeshwa

15

Jumla ya kushindwa

1204

Modeli iliyoathirika zaidi

Kimi K2.5 5
Nafasi Modeli Kampuni Idadi ya Jibu lisilo sahihi Alama Majaribio sahihi Muda wa majibu (wastani)
#154 Qwen3.5-9B none Qwen 14 4.6 4/21 1.89s
#91 GPT-5.5 none OpenAI 11 6.4 10/21 1.89s
#61 Gemini 3.1 Flash Lite low Google 9 7.2 12/21 1.89s
#123 MiMo-V2.5-Pro none Xiaomi 11 5.5 6/21 1.78s
#147 GPT-4o-mini none OpenAI 15 4.8 5/21 1.77s
#115 Qwen3.5-27B none Qwen 12 5.7 7/21 1.68s
#48 Gemini 3 Flash Preview none Google 8 7.4 13/21 1.65s
#157 Grok 4.1 Fast none X AI 13 4.4 3/19 1.62s
#128 Qwen3.6 Flash none Qwen 12 5.4 7/21 1.60s
#32 Gemini 3.5 Flash minimal Google 5 7.7 14/21 1.57s
#148 GPT-5.4 Nano none OpenAI 15 4.7 4/21 1.48s
#125 GPT-5.4 none OpenAI 13 5.5 7/21 1.42s
#87 Gemini 3.1 Flash Lite minimal Google 8 6.4 10/21 1.33s
#34 Qwen3.7 Max none Qwen 7 7.7 14/21 1.30s
#136 Elephant Alpha medium Openrouter 9 5.1 6/21 1.27s

Modeli bora kwa Idadi ya Jibu lisilo sahihi

Idadi ya Jibu lisilo sahihi dhidi ya Alama

Modeli bora kwa Muda wa majibu (wastani)