Kategori AI BENCHY
Peringkat Spesifik domain
Lihat model AI mana yang paling baik di Spesifik domain, mana yang tetap andal, dan di mana kesenjangan terbesar muncul.
Model yang ditampilkan
15
Rata-rata Skor Spesifik domain
4.8
Model terbaik
Gemini 3 Flash Preview 10.0| Peringkat | Model | Perusahaan | Skor Spesifik domain | Skor | Tes benar | Waktu respons (rata-rata) |
|---|---|---|---|---|---|---|
| #1 | Gemini 3 Flash Preview medium | 10.0 | 10.0 | 3/3 | 21.1s | |
| #2 | Gemini 3.1 Pro Preview medium | 7.7 | 9.6 | 2/3 | 32.7s | |
| #3 | Claude Opus 4.7 medium | Anthropic | 7.7 | 9.2 | 2/3 | 1.17s |
| #4 | Claude Opus 4.7 none | Anthropic | 7.7 | 9.2 | 2/3 | 1.19s |
| #14 | Gemma 4 31B medium | 7.7 | 8.3 | 2/3 | 38.5s | |
| #21 | Gemini 3 Flash Preview none | 7.7 | 8.1 | 2/3 | 963ms | |
| #42 | Claude Sonnet 4.6 none | Anthropic | 7.7 | 7.4 | 2/3 | 3.54s |
| #48 | Gemma 4 31B none | 7.7 | 6.9 | 2/3 | 3.22s | |
| #59 | Qwen3.5-Flash none | Qwen | 7.7 | 6.2 | 2/3 | 905ms |
| #63 | Qwen3.5-35B-A3B none | Qwen | 7.7 | 6.1 | 2/3 | 485ms |
| #74 | GLM 4.7 Flash none | Z.ai | 7.7 | 5.6 | 2/3 | 744ms |
| #6 | Seed-2.0-Lite medium | Bytedance Seed | 5.9 | 8.6 | 1/3 | 88.7s |
| #7 | GPT-5.3-Codex medium | OpenAI | 5.9 | 8.6 | 1/3 | 64.3s |
| #15 | Gemini 2.5 Flash medium | 5.9 | 8.2 | 1/3 | 37.3s | |
| #38 | GPT-5.4 Nano medium | OpenAI | 5.9 | 7.6 | 1/3 | 38.2s |