Navigate
AI BENCHY
Advertise here

Qwen3.7 Plus vs GLM 5.3 Flash (high)

The average score is effectively tied at 7.3 vs 7.3. GLM 5.3 Flash (high) has the lower benchmark cost at $0.029 vs $0.106. Qwen3.7 Plus is faster at 12.11s vs 25.36s, with pass rates of 54.6% vs 66.7%.

Last updated at: 2026-08-26

Rank
#104
Total Output Tokens
58,097
Response Time (avg)
12.11s
Total Cost
$0.106
Rank
#102
Total Output Tokens
78,318
Response Time (avg)
25.36s
Total Cost
$0.029
Recommended model GLM 5.3 Flash (high)

It has the best score here (7.3), while costing about 3.8x less than Qwen3.7 Plus.

Detailed comparison

Metric Qwen3.7 Plus Qwen3.7 Plus none Release: 2026-06-03 GLM 5.3 Flash GLM 5.3 Flash high Release: 2026-08-26
Score 7.3 7.3
Rank #104 #102
Reliability 10.0 10.0
Consistency 10.0 8.9
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 54.6% 66.7%
Flaky tests 0 3
Total Runs 66 66
Cost per result 0.930 0.216
Total Cost $0.106 $0.029
Input Price $0.320 / 1M $0.075 / 1M
Output Price $1.280 / 1M $0.250 / 1M
Total Input Tokens 98,833 113,003
Output Tokens 58,097 11,558
Reasoning Tokens 0 66,760
Response Time (avg) 12.11s 25.36s
Response Time (max) 206.03s 163.41s
Response Time (total) 266.40s 558.01s
Parameters ~397B total (~17B active) 320B total (18B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#104 Qwen3.7 Plus

none
Cost
$0.019
Time
213.5s
Tokens
11,960 tok

#102 GLM 5.3 Flash

high
Cost
$0.003
Time
186.7s
Tokens
10,486 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.7 Plus 5.5 10.0 33.3% 0 2.15s 7,911 639 0
GLM 5.3 Flash 8.2 7.2 88.9% 1 19.62s 7,317 373 9,973

Quick Compare

Switch Comparison Pair