Navigate
Advertise here

Mercury 2 (medium) vs LongCat 2.0 (low)

The average score is effectively tied at 6.6 vs 6.6. Mercury 2 (medium) has the lower benchmark cost at $0.121 vs $0.495. Mercury 2 (medium) is faster at 4.40s vs 115.00s, with pass rates of 53.6% vs 55.1%.

Last updated at: 2026-10-01

Compared models

Rank
#189
Total Output Tokens
96,943
Response Time (avg)
4.40s
Total Cost
$0.121
Rank
#188
Total Output Tokens
346,225
Response Time (avg)
115.00s
Total Cost
$0.495
Recommended model Mercury 2 (medium)

It has the best score here (6.6), while costing about 4.1x less than LongCat 2.0 (low).

Detailed comparison

Metric Mercury 2 Mercury 2 medium Release: 2026-02-24 LongCat 2.0 LongCat 2.0 low Release: 2026-07-20
Score 6.6 6.6
Rank #189 #188
Reliability 10.0 9.9
Consistency 8.1 8.3
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 53.6% 55.1%
Flaky tests 5 5
Total Runs 69 69
Cost per result 1.201 4.942
Total Cost $0.121 $0.495
Input Price $0.250 / 1M $0.300 / 1M
Output Price $0.750 / 1M $1.200 / 1M
Total Input Tokens 189,300 262,387
Output Tokens 10,854 7,344
Reasoning Tokens 86,089 338,881
Response Time (avg) 4.40s 115.00s
Response Time (max) 34.92s 560.39s
Response Time (total) 96.81s 2645.11s
Parameters ~100B 1.6T total (48B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#189 Mercury 2

medium
Cost
$0.002
Time
2.1s
Tokens
1,702 tok

#188 LongCat 2.0

low
Cost
$0.024
Time
428.0s
Tokens
19,557 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 8.2 7.7 77.8% 1 2.04s 7,065 296 11,328
LongCat 2.0 6.6 4.6 77.8% 2 479.30s 6,544 461 206,627

Quick Compare

Switch Comparison Pair