Navigate
Advertise here

LongCat 2.0 vs GLM 5.3 FlashX (low)

LongCat 2.0 leads on average score with 6.4 vs 6.3. LongCat 2.0 has the lower benchmark cost at $0.044 vs $0.090. LongCat 2.0 is faster at 5.17s vs 6.58s, with pass rates of 40.9% vs 65.2%.

Last updated at: 2026-09-21

Compared models

Rank
#195
Total Output Tokens
9,375
Response Time (avg)
5.17s
Total Cost
$0.044
Rank
#201
Total Output Tokens
37,036
Response Time (avg)
6.58s
Total Cost
$0.090
Recommended model LongCat 2.0

It has the best score here (6.4), while costing about 2.0x less than GLM 5.3 FlashX (low).

Detailed comparison

Metric LongCat 2.0 LongCat 2.0 none Release: 2026-07-20 GLM 5.3 FlashX GLM 5.3 FlashX low Release: 2026-09-21
Score 6.4 6.3
Rank #195 #201
Reliability 10.0 10.0
Consistency 9.3 7.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 40.9% 65.2%
Flaky tests 2 8
Total Runs 66 66
Cost per result 0.549 0.816
Total Cost $0.044 $0.090
Input Price $0.300 / 1M $0.370 / 1M
Output Price $1.200 / 1M $1.250 / 1M
Total Input Tokens 108,752 117,325
Output Tokens 9,375 10,494
Reasoning Tokens 0 26,542
Response Time (avg) 5.17s 6.58s
Response Time (max) 48.38s 73.01s
Response Time (total) 113.83s 144.85s
Parameters 1.6T total (48B active) 320B total (18B active)
Availability Open source Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#195 LongCat 2.0

none
Cost
$0.027
Time
411.7s
Tokens
22,283 tok

#201 GLM 5.3 FlashX

low
Cost
$0.002
Time
8.4s
Tokens
1,526 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
LongCat 2.0 5.5 10.0 33.3% 0 2.85s 7,446 578 0
GLM 5.3 FlashX 5.5 7.4 44.4% 1 5.45s 7,317 382 4,760

Quick Compare

Switch Comparison Pair