Kategoria ya AI BENCHY
Orodha ya Mbinu za kupinga AI
Ona ni modeli gani za AI zinafanya vizuri zaidi katika Mbinu za kupinga AI, zipi zinabaki thabiti, na pengo kubwa liko wapi. Panga kwa: Majaribio sahihi ↓.
Modeli zilizoonyeshwa
15
Wastani wa Alama ya Mbinu za kupinga AI
6.7
Modeli bora
Gemini 3 Flash Preview 10.0| Nafasi | Modeli | Kampuni | Alama ya Mbinu za kupinga AI | Alama | Majaribio sahihi | Muda wa majibu (wastani) |
|---|---|---|---|---|---|---|
| #25 | Grok 4.20 Beta medium | X AI | 8.7 | 8.0 | 3/4 | 3.16s |
| #27 | DeepSeek V3.2 medium | DeepSeek | 8.4 | 8.0 | 3/4 | 30.7s |
| #28 | GPT-5.2 Chat none | OpenAI | 8.7 | 7.9 | 3/4 | 3.40s |
| #38 | GPT-5.4 Nano medium | OpenAI | 8.3 | 7.6 | 3/4 | 4.52s |
| #41 | MiMo-V2-Flash medium | Xiaomi | 8.1 | 7.5 | 3/4 | 15.8s |
| #44 | GPT-5.4 Mini medium | OpenAI | 8.6 | 7.3 | 3/4 | 4.05s |
| #47 | Grok 4.20 medium | X AI | 8.2 | 7.0 | 3/4 | 3.36s |
| #52 | Grok 4.1 Fast medium | X AI | 8.7 | 6.7 | 3/4 | 3.81s |
| #60 | Gemma 4 26B A4B none | 8.3 | 6.2 | 3/4 | 1.28s | |
| #26 | Claude Sonnet 4.6 medium | Anthropic | 6.5 | 8.0 | 2/4 | 2.98s |
| #29 | Gemini 3.1 Flash Lite Preview none | 7.5 | 7.9 | 2/4 | 1.04s | |
| #31 | GLM 5V Turbo medium | Z.ai | 7.2 | 7.8 | 2/4 | 10.8s |
| #34 | Kimi K2.6 medium | Moonshot AI | 7.0 | 7.7 | 2/4 | 11.6s |
| #36 | GPT-5.3 Chat none | OpenAI | 6.7 | 7.7 | 2/4 | 3.86s |
| #37 | Claude Opus 4.6 medium | Anthropic | 6.4 | 7.6 | 2/4 | 7.45s |