Kategori AI BENCHY
Peringkat Spesifik domain
Lihat model AI mana yang paling baik di Spesifik domain, mana yang tetap andal, dan di mana kesenjangan terbesar muncul.
Model yang ditampilkan
15
Rata-rata Skor Spesifik domain
4.8
Model terbaik
Gemini 3 Flash Preview 10.0| Peringkat | Model | Perusahaan | Skor Spesifik domain | Skor | Tes benar | Waktu respons (rata-rata) |
|---|---|---|---|---|---|---|
| #108 | Qwen3.5-Flash none | Qwen | 7.7 | 5.8 | 2/3 | 905ms |
| #117 | Qwen3.5-35B-A3B none | Qwen | 7.7 | 5.6 | 2/3 | 485ms |
| #118 | Qwen3.6 27B none | Qwen | 7.7 | 5.6 | 2/3 | 3.03s |
| #122 | GLM 4.7 Flash none | Z.ai | 7.7 | 5.5 | 2/3 | 744ms |
| #2 | Gemini 3.5 Flash high | 7.6 | 9.6 | 2/3 | 14.1s | |
| #20 | Gemini 3.5 Flash none | 7.6 | 8.1 | 2/3 | 10.6s | |
| #5 | Qwen3.7 Max medium | Qwen | 5.9 | 9.1 | 1/3 | 24.9s |
| #15 | GPT-5.3-Codex medium | OpenAI | 5.9 | 8.4 | 1/3 | 64.3s |
| #19 | Seed-2.0-Lite medium | Bytedance Seed | 5.9 | 8.2 | 1/3 | 88.7s |
| #28 | Gemini 2.5 Flash medium | 5.9 | 7.8 | 1/3 | 37.3s | |
| #42 | GPT-5.2 medium | OpenAI | 5.9 | 7.5 | 1/3 | 77.8s |
| #64 | MiMo-V2-Flash medium | Xiaomi | 5.9 | 7.2 | 1/3 | 96.0s |
| #70 | GPT-5.4 Nano medium | OpenAI | 5.9 | 7.0 | 1/3 | 38.2s |
| #89 | Hy3 preview low | Tencent | 5.9 | 6.4 | 1/3 | 40.4s |
| #97 | Gemini 2.5 Flash none | 5.9 | 6.2 | 1/3 | 495ms |