Navigate
AI BENCHY
Advertise here

Qwen: Qwen3.7 Plus vs xAI: Grok 4.5

Grok 4.5 (medium) leads on average score with 8.3 vs 7.9. Qwen3.7 Plus (medium) has the lower benchmark cost at $0.267 vs $1.928. Qwen3.7 Plus (medium) is faster at 51.51s vs 61.71s, with pass rates of 75.8% vs 78.8%.

Recommended modelQwen3.7 Plus (medium)Its score stays close to the best score here (7.9 vs 8.3), while costing about 7.2x less than Grok 4.5 (medium).

Last updated at: 2026-07-25

Metric Qwen3.7 Plus Qwen3.7 Plus medium Release: 2026-06-03 Grok 4.5 Grok 4.5 medium Release: 2026-07-08
Score 7.9 8.3
Rank #43 #28
Reliability 10.0 10.0
Consistency 8.9 8.9
Tests Correct
Attempt pass rate 75.8% 78.8%
Flaky tests 3 3
Total Runs 66 66
Cost per result 2.072 12.049
Total Cost $0.267 $1.928
Input Price $0.320 / 1M $2.000 / 1M
Output Price $1.280 / 1M $6.000 / 1M
Total Input Tokens 115,233 122,146
Output Tokens 6,162 5,514
Reasoning Tokens 173,267 275,053
Response Time (avg) 51.51s 61.71s
Response Time (max) 315.30s 436.38s
Response Time (total) 1133.15s 1357.56s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#43 Qwen3.7 Plus

medium
Cost
$0.018
Time
193.2s
Tokens
10,821 tok

#28 xAI: Grok 4.5

medium
Cost
$0.044
Time
59.4s
Tokens
7,512 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.7 Plus 6.1 6.6 55.6% 1 108.60s 6,472 414 43,576
Grok 4.5 7.6 7.2 77.8% 1 155.69s 9,579 390 104,634

Quick Compare

Switch Comparison Pair