Navigate
Advertise here

Mercury 2.5 Preview (high) vs GLM 5.3 (low)

The average score is effectively tied at 6.5 vs 6.5. Mercury 2.5 Preview (high) has the lower benchmark cost at $0.039 vs $0.473. Mercury 2.5 Preview (high) is faster at 7.61s vs 15.28s, with pass rates of 65.2% vs 63.8%.

Last updated at: 2026-10-01

Compared models

Rank
#193
Total Output Tokens
180,597
Response Time (avg)
7.61s
Total Cost
$0.039
Rank
#195
Total Output Tokens
28,602
Response Time (avg)
15.28s
Total Cost
$0.473
Recommended model Mercury 2.5 Preview (high)

It has the best score here (6.5), while costing about 12.4x less than GLM 5.3 (low).

Detailed comparison

Metric Mercury 2.5 Preview Mercury 2.5 Preview high Release: 2026-09-02 GLM 5.3 GLM 5.3 low Release: 2026-08-20
Score 6.5 6.5
Rank #193 #195
Reliability 9.9 10.0
Consistency 8.0 8.3
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 65.2% 63.8%
Flaky tests 6 5
Total Runs 69 69
Cost per result 0.319 3.936
Total Cost $0.039 $0.473
Input Price $0.000 / 1M $1.400 / 1M
Output Price $0.000 / 1M $4.400 / 1M
Total Input Tokens 298,512 247,436
Output Tokens 3,178 11,005
Reasoning Tokens 177,419 17,597
Response Time (avg) 7.61s 15.28s
Response Time (max) 86.84s 83.60s
Response Time (total) 175.08s 351.49s
Parameters ~100B 744B total (40B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#193 Mercury 2.5 Preview

high
Cost
$0.001
Time
4.5s
Tokens
3,851 tok

#195 GLM 5.3

low
Cost
$0.007
Time
28.4s
Tokens
1,599 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 Preview 6.7 7.8 55.6% 1 6.13s 7,751 485 42,569
GLM 5.3 6.2 6.9 55.6% 1 21.02s 7,317 368 6,764

Quick Compare

Switch Comparison Pair