Navigate
Advertise here

Mercury 2.5 (high) vs GLM 5.3 (low)

The average score is effectively tied at 6.5 vs 6.5. Mercury 2.5 (high) has the lower benchmark cost at $0.044 vs $0.473. Mercury 2.5 (high) is faster at 7.62s vs 15.28s, with pass rates of 59.4% vs 63.8%.

Last updated at: 2026-10-01

Compared models

Rank
#194
Total Output Tokens
193,676
Response Time (avg)
7.62s
Total Cost
$0.044
Rank
#195
Total Output Tokens
28,602
Response Time (avg)
15.28s
Total Cost
$0.473
Recommended model Mercury 2.5 (high)

It has the best score here (6.5), while costing about 10.9x less than GLM 5.3 (low).

Detailed comparison

Metric Mercury 2.5 Mercury 2.5 high Release: 2026-09-08 GLM 5.3 GLM 5.3 low Release: 2026-08-20
Score 6.5 6.5
Rank #194 #195
Reliability 9.8 10.0
Consistency 8.7 8.3
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 59.4% 63.8%
Flaky tests 4 5
Total Runs 69 69
Cost per result 0.362 3.936
Total Cost $0.044 $0.473
Input Price $0.040 / 1M $1.400 / 1M
Output Price $0.150 / 1M $4.400 / 1M
Total Input Tokens 358,562 247,436
Output Tokens 4,266 11,005
Reasoning Tokens 189,410 17,597
Response Time (avg) 7.62s 15.28s
Response Time (max) 80.30s 83.60s
Response Time (total) 175.36s 351.49s
Parameters ~100B 744B total (40B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#194 Mercury 2.5

high
Cost
$0.001
Time
3.5s
Tokens
2,447 tok

#195 GLM 5.3

low
Cost
$0.007
Time
28.4s
Tokens
1,599 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 6.4 7.8 44.4% 1 6.58s 7,757 415 44,012
GLM 5.3 6.2 6.9 55.6% 1 21.02s 7,317 368 6,764

Quick Compare

Switch Comparison Pair