AI BENCHY
Your ad here

AI BENCHY Failures

Wrong answer Failures

See which AI models run into Wrong answer most often, so you can spot reliability risks before choosing one. Sort by: Response Time (avg) ↑.

Models Shown

15

Total Failures

572

Most Affected Model

Mercury 2 13
Rank Model Company Wrong answer Count Score Tests Correct Response Time (avg)
#95 Grok 4.1 Fast none X AI 13 4.5 3/18 1.76s
#55 MiMo-V2-Omni none Xiaomi 8 6.5 8/18 1.99s
#89 GPT-4o-mini none OpenAI 13 4.9 4/18 2.00s
#69 Kimi K2.6 none Moonshot AI 8 5.8 7/18 2.05s
#54 Mercury 2 medium Inception 6 6.5 8/18 2.21s
#65 MiMo-V2-Pro none Xiaomi 9 6.0 7/18 2.39s
#61 Seed-2.0-Lite none Bytedance Seed 10 6.2 8/18 2.53s
#49 Qwen3.5 Plus 2026-02-15 none Qwen 9 6.8 9/18 2.60s
#94 MiMo-V2-Flash none Xiaomi 12 4.5 3/18 2.79s
#77 GLM 5 Turbo none Z.ai 10 5.5 6/18 2.94s
#58 GLM 5V Turbo none Z.ai 8 6.2 8/18 3.10s
#4 Claude Opus 4.7 none Anthropic 2 9.2 16/18 3.13s
#22 Gemini 3.1 Flash Lite Preview low Google 4 8.1 13/18 3.22s
#59 Qwen3.5-Flash none Qwen 9 6.2 8/18 3.25s
#74 GLM 4.7 Flash none Z.ai 10 5.6 5/18 3.35s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)