Kategori AI BENCHY
Peringkat Spesifik domain
Lihat model AI mana yang paling baik di Spesifik domain, mana yang tetap andal, dan di mana kesenjangan terbesar muncul. Urutkan berdasarkan: Tes benar ↓.
Model yang ditampilkan
15
Rata-rata Skor Spesifik domain
4.8
Model terbaik
Gemini 3 Flash Preview 10.0| Peringkat | Model | Perusahaan | Skor Spesifik domain | Skor | Tes benar | Waktu respons (rata-rata) |
|---|---|---|---|---|---|---|
| #107 | Laguna Xs.2 medium | Poolside | 4.1 | 5.8 | 0/3 | 11.1s |
| #110 | Seed-2.0-Lite none | Bytedance Seed | 3.6 | 5.8 | 0/3 | 1.33s |
| #112 | GLM 5.1 none | Z.ai | 2.9 | 5.7 | 0/3 | 1.99s |
| #115 | Qwen3.5-27B none | Qwen | 3.0 | 5.7 | 0/3 | 540ms |
| #119 | Cobuddy medium | Baidu | 2.9 | 5.6 | 0/3 | 128.2s |
| #126 | gpt-oss-120b none | OpenAI | 3.0 | 5.4 | 0/3 | 35.0s |
| #127 | Grok 4.20 none | X AI | 3.0 | 5.4 | 0/3 | 687ms |
| #129 | MiniMax M2.5 medium | Minimax | 2.9 | 5.3 | 0/3 | 237.3s |
| #130 | MiniMax M2.7 medium | Minimax | 3.0 | 5.3 | 0/3 | 19.0s |
| #133 | DeepSeek V3.2 none | DeepSeek | 2.9 | 5.2 | 0/3 | 4.17s |
| #136 | Elephant Alpha medium | Openrouter | 3.0 | 5.1 | 0/3 | 925ms |
| #137 | Elephant Alpha none | Openrouter | 3.0 | 5.1 | 0/3 | 927ms |
| #138 | Ling-2.6-flash none | Inclusionai | 3.0 | 5.0 | 0/3 | 4.95s |
| #141 | Nemotron 3 Super none | NVIDIA | 3.6 | 4.9 | 0/3 | 6.23s |
| #143 | MiMo-V2.5 none | Xiaomi | 3.0 | 4.9 | 0/3 | 756ms |