AI BENCHY
Advertise here

AI BENCHY Category Failures

Trivia: Wrong answer

Trivia
Wrong answer

See which AI models are most likely to hit Wrong answer on Trivia, so you can spot weak points faster.

Models Shown

15

Total Failures

117

Most Affected Model

Claude Opus 4.7 1

Failure Reasons

Rank Model Company Wrong answer Count Category Score Tests Correct Response Time (avg)
#56 Seed-2.0-Mini medium Bytedance Seed 1 3.0 0/1 56.8s
#57 Qwen3.5-35B-A3B medium Qwen 1 3.0 0/1 177.4s
#58 GPT-5.2 medium OpenAI 1 3.0 0/1 28.2s
#59 DeepSeek V3.2 medium DeepSeek 1 3.0 0/1 84.0s
#60 GPT-5.4 Mini medium OpenAI 1 3.0 0/1 30.1s
#61 Claude Sonnet 4.6 none Anthropic 1 3.0 0/1 4.67s
#62 MiMo-V2-Omni medium Xiaomi 1 3.0 0/1 234.2s
#64 Gemma 4 31B none Google 1 3.0 0/1 1.25s
#65 DeepSeek V4 Pro high DeepSeek 1 3.0 0/1 39.1s
#66 Grok 4.20 medium X AI 1 3.0 0/1 63.5s
#67 GPT-5 Mini medium OpenAI 1 3.0 0/1 9.99s
#68 Gemini 3.1 Flash Lite minimal Google 1 3.0 0/1 724ms
#69 Kimi K2.5 medium Moonshot AI 1 3.0 0/1 83.9s
#70 Qwen3.6 27B medium Qwen 1 3.0 0/1 81.0s
#72 GPT-5.5 none OpenAI 1 3.0 0/1 5.01s

Top Models by Wrong answer Count

Wrong answer Count vs Score

Top Models by Response Time (avg)

Top Models by Estimated Wasted Cost