Kategori AI BENCHY
Peringkat Pengetahuan umum
Lihat model AI mana yang paling baik di Pengetahuan umum, mana yang tetap andal, dan di mana kesenjangan terbesar muncul. Urutkan berdasarkan: Metrik ↑.
Model yang ditampilkan
11
Rata-rata Skor Pengetahuan umum
2.9
Model terbaik
Gemini 3 PRO Preview 0.0Alasan kegagalan
| Peringkat | Model | Perusahaan | Skor Pengetahuan umum | Skor | Tes benar | Waktu respons (rata-rata) |
|---|---|---|---|---|---|---|
| #132 | Qwen3.5-9B none | Qwen | 3.0 | 4.7 | 0/1 | 2.32s |
| #133 | HY3 Preview none | Tencent | 3.0 | 4.6 | 0/1 | 2.71s |
| #135 | GPT-5.4 Nano none | OpenAI | 3.0 | 4.5 | 0/1 | 773ms |
| #136 | GLM 4.7 Flash medium | Z.ai | 3.0 | 4.5 | 0/1 | 11.1s |
| #137 | MiMo-V2-Flash none | Xiaomi | 3.0 | 4.5 | 0/1 | 1.82s |
| #139 | Grok 4.1 Fast none | X AI | 3.0 | 4.4 | 0/1 | 731ms |
| #140 | Qwen3.5-9B medium | Qwen | 3.0 | 4.3 | 0/1 | 177.0s |
| #142 | Granite 4.1 8B none | IBM Granite | 3.0 | 4.1 | 0/1 | 306ms |
| #1 | Gemini 3 Flash Preview medium | 10.0 | 10.0 | 1/1 | 5.50s | |
| #2 | Gemini 3.1 Pro Preview medium | 10.0 | 9.6 | 1/1 | 6.27s | |
| #7 | Gemini 3 Flash Preview low | 10.0 | 8.8 | 1/1 | 2.75s |