Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Qwen: Qwen3.5-122B-A10B vs xAI: Grok Build 0.1

Last updated at: 2026-05-21

Metric Qwen3.5-122B-A10B Qwen3.5-122B-A10B medium Release: 2026-02-24 Grok Build 0.1 Grok Build 0.1 medium Release: 2026-05-21
Score 7.9 7.8
Rank #36 #41
Reliability 10.0 10.0
Consistency 8.7 8.9
Tests Correct
Attempt pass rate 75.4% 71.9%
Flaky tests 3 3
Total Runs 57 57
Cost per result 4.315 4.064
Total Cost $0.561 $0.488
Input Price $0.260 / 1M $1.000 / 1M
Output Price $2.080 / 1M $2.000 / 1M
Output Tokens 18,457 1,947
Reasoning Tokens 177,734 223,372
Response Time (avg) 32.51s 22.28s
Response Time (max) 119.29s 88.28s
Response Time (total) 617.70s 423.30s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 10.0 10.0 100.0% 0 9.75s 269 16,835
Grok Build 0.1 10.0 10.0 100.0% 0 5.46s 195 9,825
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 4.7 1.6 66.7% 1 70.98s 322 10,694
Grok Build 0.1 7.3 3.7 66.7% 1 30.98s 354 17,734
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 10.0 10.0 100.0% 0 107.79s 483 11,337
Grok Build 0.1 10.0 10.0 100.0% 0 30.81s 231 18,779
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 10.0 10.0 100.0% 0 23.41s 270 16,558
Grok Build 0.1 10.0 10.0 100.0% 0 7.76s 180 10,343
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 2.9 7.2 11.1% 1 63.40s 15,537 64,889
Grok Build 0.1 5.3 10.0 33.3% 0 77.75s 501 111,807
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 3.4 2.2 33.3% 1 34.11s 66 7,592
Grok Build 0.1 3.8 2.5 33.3% 1 10.14s 78 5,386
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 10.0 10.0 100.0% 0 9.88s 77 7,372
Grok Build 0.1 9.8 10.0 100.0% 0 9.62s 57 12,436
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 10.0 10.0 100.0% 0 17.18s 289 26,165
Grok Build 0.1 6.2 7.5 55.6% 1 8.67s 161 15,476
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 10.0 10.0 100.0% 0 4.60s 322 1,226
Grok Build 0.1 10.0 10.0 100.0% 0 9.40s 180 5,319
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 3.0 10.0 0.0% 0 52.87s 822 15,066
Grok Build 0.1 3.0 10.0 0.0% 0 26.07s 10 16,267

Quick Compare

Switch Comparison Pair