Navigate
AI BENCHY
Advertise here

Inception: Mercury 2 vs Kwaipilot: KAT-Coder-Pro V2.5

KAT-Coder-Pro V2.5 (high) leads on average score with 7.2 vs 7.0. Mercury 2 (medium) has the lower benchmark cost at $0.093 vs $0.482. Mercury 2 (medium) is faster at 2.72s vs 20.83s, with pass rates of 51.5% vs 63.6%.

Recommended modelMercury 2 (medium)Its score stays close to the best score here (7.0 vs 7.2), while costing about 5.2x less than KAT-Coder-Pro V2.5 (high).

Last updated at: 2026-07-18

Metric Mercury 2 Mercury 2 medium Release: 2026-02-24 KAT-Coder-Pro V2.5 KAT-Coder-Pro V2.5 high Release: 2026-07-14
Score 7.0 7.2
Rank #77 #68
Reliability 10.0 10.0
Consistency 8.8 7.8
Tests Correct
Attempt pass rate 51.5% 63.6%
Flaky tests 3 6
Total Runs 66 66
Cost per result 0.928 4.378
Total Cost $0.093 $0.482
Input Price $0.250 / 1M $0.740 / 1M
Output Price $0.750 / 1M $2.960 / 1M
Total Input Tokens 109,572 106,076
Output Tokens 10,313 9,071
Reasoning Tokens 76,806 127,093
Response Time (avg) 2.72s 20.83s
Response Time (max) 14.63s 199.97s
Response Time (total) 57.12s 458.31s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#77 Mercury 2

medium
Cost
$0.002
Time
2.1s
Tokens
1,702 tok

#68 KAT-Coder-Pro V2.5

high
Cost
$0.009
Time
26.4s
Tokens
3,117 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 8.2 7.7 77.8% 1 2.04s 7,065 296 11,328
KAT-Coder-Pro V2.5 6.4 7.9 44.4% 1 22.00s 7,893 422 20,461

Quick Compare

Switch Comparison Pair