Kategori AI BENCHY
Peringkat Pengetahuan umum
Lihat model AI mana yang paling baik di Pengetahuan umum, mana yang tetap andal, dan di mana kesenjangan terbesar muncul.
Model yang ditampilkan
15
Rata-rata Skor Pengetahuan umum
2.9
Model terbaik
Gemini 3 Flash Preview 10.0Alasan kegagalan
| Peringkat | Model | Perusahaan | Skor Pengetahuan umum | Skor | Tes benar | Waktu respons (rata-rata) |
|---|---|---|---|---|---|---|
| #39 | HY3 Preview low | Tencent | 3.0 | 7.7 | 0/1 | 41.7s |
| #40 | Gemini 3.1 Flash Lite Preview none | 3.0 | 7.7 | 0/1 | 814ms | |
| #41 | GPT-5.2 Chat none | OpenAI | 3.0 | 7.6 | 0/1 | 6.89s |
| #42 | Kimi K2.6 medium | Moonshot AI | 3.0 | 7.6 | 0/1 | 130.3s |
| #43 | Step 3.5 Flash medium | Stepfun | 3.0 | 7.6 | 0/1 | 108.4s |
| #44 | Gemini 3.1 Flash Lite low | 3.0 | 7.6 | 0/1 | 1.46s | |
| #45 | Qwen3.5-Flash medium | Qwen | 3.0 | 7.6 | 0/1 | 49.0s |
| #46 | GPT-5.3 Chat none | OpenAI | 3.0 | 7.6 | 0/1 | 4.38s |
| #47 | GLM 5.1 medium | Z.ai | 3.0 | 7.6 | 0/1 | 29.4s |
| #48 | DeepSeek V4 Flash high | DeepSeek | 3.0 | 7.6 | 0/1 | 54.5s |
| #49 | GLM 5V Turbo medium | Z.ai | 3.0 | 7.5 | 0/1 | 41.0s |
| #50 | Qwen3.6 Flash medium | Qwen | 3.0 | 7.5 | 0/1 | 122.9s |
| #52 | Claude Opus 4.6 medium | Anthropic | 3.0 | 7.4 | 0/1 | 63.2s |
| #53 | GPT-5.4 Nano medium | OpenAI | 3.0 | 7.3 | 0/1 | 4.81s |
| #54 | Qwen3.6 Max Preview none | Qwen | 3.0 | 7.2 | 0/1 | 1.97s |