Navigate
AI BENCHY
Advertise here

Qwen3.5-122B-A10B (medium) vs Grok 4.20 (medium)

The average score is effectively tied at 7.1 vs 7.1. Grok 4.20 (medium) has the lower benchmark cost at $0.777 vs $1.046. Grok 4.20 (medium) is faster at 29.47s vs 64.16s, with pass rates of 71.2% vs 63.6%.

Last updated at: 2026-07-25

Rank
#80
Total Output Tokens
487,218
Response Time (avg)
64.16s
Total Cost
$1.046
Rank
#83
Total Output Tokens
259,340
Response Time (avg)
29.47s
Total Cost
$0.777
Recommended model Grok 4.20 (medium)

It has the best score here (7.1), while responding about 2.2x faster than Qwen3.5-122B-A10B (medium).

Detailed comparison

Metric Qwen3.5-122B-A10B Qwen3.5-122B-A10B medium Release: 2026-02-24 Grok 4.20 Grok 4.20 medium Release: 2026-03-31
Score 7.1 7.1
Rank #80 #83
Reliability 10.0 10.0
Consistency 8.5 8.5
Tests Correct
Attempt pass rate 71.2% 63.6%
Flaky tests 4 4
Total Runs 66 66
Cost per result 8.509 9.709
Total Cost $1.046 $0.777
Input Price $0.260 / 1M $1.250 / 1M
Output Price $2.080 / 1M $2.500 / 1M
Total Input Tokens 124,771 102,791
Output Tokens 44,077 5,363
Reasoning Tokens 443,141 253,977
Response Time (avg) 64.16s 29.47s
Response Time (max) 519.30s 199.66s
Response Time (total) 1411.60s 648.35s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#80 Qwen3.5-122B-A10B

medium
Cost
$0.019
Time
48.7s
Tokens
6,034 tok

#83 xAI: Grok 4.20

medium
Cost
$0.041
Time
110.3s
Tokens
16,336 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 6.0 7.2 55.6% 1 114.48s 7,630 8,057 82,578
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150

Quick Compare

Switch Comparison Pair