Navigate
Advertise here

Grok 4.3 (medium) vs GLM 5.3 (low)

The average score is effectively tied at 6.5 vs 6.5. GLM 5.3 (low) has the lower benchmark cost at $0.473 vs $0.989. GLM 5.3 (low) is faster at 15.28s vs 44.53s, with pass rates of 63.8% vs 63.8%.

Last updated at: 2026-10-01

Compared models

Rank
#199
Total Output Tokens
232,542
Response Time (avg)
44.53s
Total Cost
$0.989
Rank
#195
Total Output Tokens
28,602
Response Time (avg)
15.28s
Total Cost
$0.473
Recommended model GLM 5.3 (low)

It has the best score here (6.5), while costing about 2.1x less than Grok 4.3 (medium).

Detailed comparison

Metric Grok 4.3 Grok 4.3 medium Release: 2026-05-01 GLM 5.3 GLM 5.3 low Release: 2026-08-20
Score 6.5 6.5
Rank #199 #195
Reliability 10.0 10.0
Consistency 8.3 8.3
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 63.8% 63.8%
Flaky tests 5 5
Total Runs 69 69
Cost per result 8.235 3.936
Total Cost $0.989 $0.473
Input Price $1.250 / 1M $1.400 / 1M
Output Price $2.500 / 1M $4.400 / 1M
Total Input Tokens 325,471 247,436
Output Tokens 14,867 11,005
Reasoning Tokens 217,675 17,597
Response Time (avg) 44.53s 15.28s
Response Time (max) 216.69s 83.60s
Response Time (total) 1024.25s 351.49s
Parameters ~1.5T total (~150B active) 744B total (40B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#199 SpaceXAI: Grok 4.3

medium
Cost
$0.009
Time
19.0s
Tokens
3,661 tok

#195 GLM 5.3

low
Cost
$0.007
Time
28.4s
Tokens
1,599 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Grok 4.3 5.9 7.7 44.4% 1 41.23s 8,340 1,028 31,226
GLM 5.3 6.2 6.9 55.6% 1 21.02s 7,317 368 6,764

Quick Compare

Switch Comparison Pair