AI BENCHY Category
Domain specific Ranking
See which AI models perform best on Domain specific, which ones stay reliable, and where the biggest gaps appear. Sort by: Tests Correct ↓.
| Rank | Model | Company | Domain specific Score | Score | Tests Correct | Response Time (avg) |
|---|---|---|---|---|---|---|
| #21 | GPT-5.4 medium | OpenAI | 5.3 | 8.0 | 1/3 | 74.3s |
| #24 | GPT-5.2 Chat none | OpenAI | 5.3 | 7.9 | 1/3 | 17.8s |
| #25 | Qwen3.5 Plus 2026-02-15 medium | Qwen | 5.3 | 7.9 | 1/3 | 17.5s |
| #28 | Gemini 2.5 Flash medium | 5.9 | 7.8 | 1/3 | 37.3s | |
| #30 | Qwen3.5-27B medium | Qwen | 5.3 | 7.8 | 1/3 | 79.5s |
| #33 | Hy3 preview medium | Tencent | 5.3 | 7.7 | 1/3 | 22.3s |
| #35 | Gemini 3 PRO Preview medium | 5.3 | 7.6 | 1/3 | 7.01s | |
| #38 | Grok 4.3 medium | X AI | 5.3 | 7.6 | 1/3 | 181.7s |
| #42 | GPT-5.2 medium | OpenAI | 5.9 | 7.5 | 1/3 | 77.8s |
| #43 | MiMo-V2.5-Pro medium | Xiaomi | 5.3 | 7.5 | 1/3 | 37.9s |
| #46 | Qwen3.6 35B A3B medium | Qwen | 5.3 | 7.4 | 1/3 | 22.5s |
| #47 | Grok Build 0.1 medium | X AI | 5.3 | 7.4 | 1/3 | 158.0s |
| #49 | Qwen3.5-Flash medium | Qwen | 5.3 | 7.4 | 1/3 | 146.5s |
| #50 | Gemini 3.1 Flash Lite Preview low | 5.3 | 7.4 | 1/3 | 2.36s | |
| #51 | Mimo V2 PRO medium | Xiaomi | 5.3 | 7.4 | 1/3 | 8.82s |