Navigate
Advertise here

Mercury 2.5 (medium) vs KAT-Coder-Pro V2.5

The average score is effectively tied at 6.8 vs 6.7. Mercury 2.5 (medium) has the lower benchmark cost at $0.022 vs $0.487. Mercury 2.5 (medium) is faster at 3.19s vs 26.01s, with pass rates of 65.2% vs 71.2%.

Last updated at: 2026-09-08

Compared models

Rank
#157
Total Output Tokens
120,129
Response Time (avg)
3.19s
Total Cost
$0.022
Rank
#161
Total Output Tokens
139,572
Response Time (avg)
26.01s
Total Cost
$0.487
Recommended model Mercury 2.5 (medium)

It has the best score here (6.8), while costing about 22.4x less than KAT-Coder-Pro V2.5.

Detailed comparison

Metric Mercury 2.5 Mercury 2.5 medium Release: 2026-09-08 KAT-Coder-Pro V2.5 KAT-Coder-Pro V2.5 none Release: 2026-07-14
Score 6.8 6.7
Rank #157 #161
Reliability 9.8 10.0
Consistency 8.6 7.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 65.2% 71.2%
Flaky tests 4 8
Total Runs 66 66
Cost per result 0.181 4.419
Total Cost $0.022 $0.487
Input Price $0.040 / 1M $0.740 / 1M
Output Price $0.150 / 1M $2.960 / 1M
Total Input Tokens 91,402 98,508
Output Tokens 3,138 139,572
Reasoning Tokens 116,991 0
Response Time (avg) 3.19s 26.01s
Response Time (max) 9.87s 335.41s
Response Time (total) 70.18s 572.22s
Parameters ~100B ~700B total (~72B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#157 Mercury 2.5

medium
Cost
$0.001
Time
3.8s
Tokens
3,386 tok

#161 KAT-Coder-Pro V2.5

none
Cost
$0.010
Time
29.0s
Tokens
3,439 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 5.0 5.1 44.5% 2 3.70s 7,807 488 21,857
KAT-Coder-Pro V2.5 6.1 4.7 66.7% 2 22.52s 7,893 22,440 0

Quick Compare

Switch Comparison Pair