Navigate
Advertise here

Qwen3.7 Max vs GLM 5.3 FlashX (max)

The average score is effectively tied at 7.9 vs 7.9. Qwen3.7 Max has the lower benchmark cost at $0.492 vs $0.623. Qwen3.7 Max is faster at 7.85s vs 43.97s, with pass rates of 69.6% vs 78.3%.

Last updated at: 2026-10-01

Compared models

Rank
#100
Total Output Tokens
20,670
Response Time (avg)
7.85s
Total Cost
$0.492
Rank
#97
Total Output Tokens
402,765
Response Time (avg)
43.97s
Total Cost
$0.623
Recommended model Qwen3.7 Max

It has the best score here (7.9), while responding about 5.6x faster than GLM 5.3 FlashX (max).

Detailed comparison

Metric Qwen3.7 Max Qwen3.7 Max none Release: 2026-05-22 GLM 5.3 FlashX GLM 5.3 FlashX max Release: 2026-09-21
Score 7.9 7.9
Rank #100 #97
Reliability 9.9 10.0
Consistency 10.0 8.0
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 69.6% 78.3%
Flaky tests 0 6
Total Runs 69 69
Cost per result 3.070 4.153
Total Cost $0.492 $0.623
Input Price $1.475 / 1M $0.370 / 1M
Output Price $4.425 / 1M $1.250 / 1M
Total Input Tokens 243,553 322,894
Output Tokens 20,670 8,010
Reasoning Tokens 0 394,755
Response Time (avg) 7.85s 43.97s
Response Time (max) 80.34s 390.70s
Response Time (total) 180.65s 1011.29s
Parameters ~1T total (~40B active) 320B total (18B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#100 Qwen3.7 Max

none
Cost
$0.046
Time
195.0s
Tokens
12,171 tok

#97 GLM 5.3 FlashX

max
Cost
$0.046
Time
251.3s
Tokens
36,769 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.7 Max 5.5 10.0 33.3% 0 1.35s 7,911 582 0
GLM 5.3 FlashX 8.2 7.2 88.9% 1 24.61s 7,317 463 36,503

Quick Compare

Switch Comparison Pair