AI BENCHY
Advertise here

AI BENCHY Category Failures

Domain specific: Wrong answer

Domain specific
Wrong answer

See which AI models are most likely to hit Wrong answer on Domain specific, so you can spot weak points faster. Sort by: Tests Correct ↓.

Models Shown

15

Total Failures

314

Most Affected Model

Gemini 3.5 Flash 1
Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#135 Kimi K2.5 none Moonshot AI 2 5.3 1/3 4.38s
#139 DeepSeek V4 Flash none DeepSeek 2 5.3 1/3 19.7s
#140 Qwen3 Coder Next none Qwen 2 5.3 1/3 962ms
#142 Mistral Small 4 none Mistral 2 5.3 1/3 367ms
#146 Laguna Xs.2 none Poolside 2 5.3 1/3 371ms
#150 Qwen3 Coder Next medium Qwen 2 5.3 1/3 638ms
#151 Trinity Large Preview none Arcee AI 2 5.3 1/3 877ms
#152 MiMo-V2-Flash none Xiaomi 2 5.3 1/3 564ms
#155 Mercury 2 none Inception 2 5.3 1/3 534ms
#157 Grok 4.1 Fast none X AI 2 5.9 1/3 1.06s
#160 LFM2-24B-A2B none Liquid 1 5.9 1/3 287ms
#14 Qwen3.6 Max Preview medium Qwen 3 2.9 0/3 95.9s
#17 GLM 5 medium Z.ai 2 3.5 0/3 0ms
#18 Qwen3.7 Plus medium Qwen 3 3.6 0/3 45.3s
#23 GLM 5 Turbo medium Z.ai 2 2.9 0/3 71.1s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost