AI BENCHY Category
Trivia Ranking
See which AI models perform best on Trivia, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↓.
Failure Reasons
| Rank | Model | Company | Trivia Score | Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #83 | GPT-5 Nano medium | OpenAI | 3.0 | 6.2 | 0/1 | 20.1s |
| #84 | DeepSeek V4 Pro none | DeepSeek | 3.0 | 6.2 | 0/1 | 15.6s |
| #85 | Nemotron 3 Super medium | NVIDIA | 3.0 | 6.1 | 0/1 | 55.3s |
| #86 | Seed-2.0-Lite none | Bytedance Seed | 3.0 | 6.0 | 0/1 | 1.96s |
| #87 | GLM 5V Turbo none | Z.ai | 3.0 | 6.0 | 0/1 | 2.23s |
| #88 | Owl Alpha medium | Openrouter | 3.0 | 6.0 | 0/1 | 2.38s |
| #89 | Qwen3.5-Flash none | Qwen | 3.0 | 6.0 | 0/1 | 588ms |
| #90 | Qwen3.5 Plus 2026-04-20 none | Qwen | 3.0 | 5.9 | 0/1 | 33.3s |
| #91 | Qwen3.5-35B-A3B none | Qwen | 3.0 | 5.9 | 0/1 | 493ms |
| #92 | MiMo-V2-Pro none | Xiaomi | 3.0 | 5.9 | 0/1 | 1.63s |
| #93 | Qwen3.5-27B none | Qwen | 3.0 | 5.9 | 0/1 | 599ms |
| #94 | Qwen3.6 27B none | Qwen | 3.0 | 5.8 | 0/1 | 4.03s |
| #95 | Cobuddy medium | Baidu | 3.0 | 5.8 | 0/1 | 37.0s |
| #96 | Owl Alpha none | Openrouter | 3.0 | 5.8 | 0/1 | 2.50s |
| #97 | GLM 4.7 Flash none | Z.ai | 3.0 | 5.8 | 0/1 | 692ms |