AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Category

Instructions following Ranking

See which AI models perform best on Instructions following, which ones stay reliable, and where the biggest gaps appear. Sort by: Response Time (avg) ↓.

Models Shown

8

Average Instructions following Score

8.0

Best Model

Kimi K2.5 10.0
Rank Model Company Instructions following Score Score Tests Correct Response Time (avg)
#86 GPT-5.4 Mini none OpenAI 6.3 5.1 1/2 728ms
#79 Grok 4.20 Beta none X AI 4.8 5.3 0/2 687ms
#62 Gemini 2.5 Flash none Google 8.0 6.2 1/2 672ms
#70 Qwen3.5-122B-A10B none Qwen 4.5 5.7 0/2 585ms
#91 Mercury 2 none Inception 6.5 4.8 1/2 551ms
#90 Qwen3.5-9B none Qwen 6.5 4.8 1/2 514ms
#82 Grok 4.20 none X AI 4.8 5.2 0/2 455ms
#83 Mistral Small 4 none Mistral 6.5 5.2 1/2 380ms

Top Models by Instructions following Score

Instructions following Score vs Total Cost

Top Models by Response Time (avg)