Navigate
Advertise here

Grok 4.3 (medium) vs GLM 5.1

The average score is effectively tied at 6.5 vs 6.5. GLM 5.1 has the lower benchmark cost at $0.311 vs $0.989. GLM 5.1 is faster at 11.10s vs 44.53s, with pass rates of 63.8% vs 47.8%.

Last updated at: 2026-10-01

Compared models

Rank
#199
Total Output Tokens
232,542
Response Time (avg)
44.53s
Total Cost
$0.989
Rank
#196
Total Output Tokens
20,076
Response Time (avg)
11.10s
Total Cost
$0.311
Recommended model GLM 5.1

It has the best score here (6.5), while costing about 3.2x less than Grok 4.3 (medium).

Detailed comparison

Metric Grok 4.3 Grok 4.3 medium Release: 2026-05-01 GLM 5.1 GLM 5.1 none Release: 2026-04-07
Score 6.5 6.5
Rank #199 #196
Reliability 10.0 10.0
Consistency 8.3 8.2
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 63.8% 47.8%
Flaky tests 5 5
Total Runs 69 69
Cost per result 8.235 3.478
Total Cost $0.989 $0.311
Input Price $1.250 / 1M $0.965 / 1M
Output Price $2.500 / 1M $3.032 / 1M
Total Input Tokens 325,471 258,956
Output Tokens 14,867 20,076
Reasoning Tokens 217,675 0
Response Time (avg) 44.53s 11.10s
Response Time (max) 216.69s 107.19s
Response Time (total) 1024.25s 255.39s
Parameters ~1.5T total (~150B active) 744B total (40B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#199 SpaceXAI: Grok 4.3

medium
Cost
$0.009
Time
19.0s
Tokens
3,661 tok

#196 GLM 5.1

none
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Grok 4.3 5.9 7.7 44.4% 1 41.23s 8,340 1,028 31,226
GLM 5.1 3.9 9.7 0.0% 0 4.96s 7,256 525 0

Quick Compare

Switch Comparison Pair