Navigate
Advertise here

Mercury 2.5 (medium) vs LongCat 2.0 (low)

Mercury 2.5 (medium) leads on average score with 6.8 vs 6.7. Mercury 2.5 (medium) has the lower benchmark cost at $0.022 vs $0.412. Mercury 2.5 (medium) is faster at 3.19s vs 113.61s, with pass rates of 65.2% vs 54.6%.

Last updated at: 2026-09-08

Compared models

Rank
#157
Total Output Tokens
120,129
Response Time (avg)
3.19s
Total Cost
$0.022
Rank
#162
Total Output Tokens
320,594
Response Time (avg)
113.61s
Total Cost
$0.412
Recommended model Mercury 2.5 (medium)

It has the best score here (6.8), while costing about 19.0x less than LongCat 2.0 (low).

Detailed comparison

Metric Mercury 2.5 Mercury 2.5 medium Release: 2026-09-08 LongCat 2.0 LongCat 2.0 low Release: 2026-07-20
Score 6.8 6.7
Rank #157 #162
Reliability 9.8 9.9
Consistency 8.6 8.5
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 65.2% 54.6%
Flaky tests 4 4
Total Runs 66 66
Cost per result 0.181 4.120
Total Cost $0.022 $0.412
Input Price $0.040 / 1M $0.300 / 1M
Output Price $0.150 / 1M $1.200 / 1M
Total Input Tokens 91,402 90,870
Output Tokens 3,138 5,676
Reasoning Tokens 116,991 314,918
Response Time (avg) 3.19s 113.61s
Response Time (max) 9.87s 560.39s
Response Time (total) 70.18s 2499.46s
Parameters ~100B 1.6T total (48B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#157 Mercury 2.5

medium
Cost
$0.001
Time
3.8s
Tokens
3,386 tok

#162 LongCat 2.0

low
Cost
$0.024
Time
428.0s
Tokens
19,557 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 5.0 5.1 44.5% 2 3.70s 7,807 488 21,857
LongCat 2.0 6.6 4.6 77.8% 2 479.30s 6,544 461 206,627

Quick Compare

Switch Comparison Pair