Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

KAT-Coder-Pro V2.5 (medium) vs Grok 4.20 (medium)

Grok 4.20 (medium) leads on average score with 7.0 vs 6.9. KAT-Coder-Pro V2.5 (medium) has the lower benchmark cost at $0.478 vs $0.805. KAT-Coder-Pro V2.5 (medium) is faster at 24.87s vs 32.56s, with pass rates of 65.2% vs 60.6%.

Last updated at: 2026-08-15

Rank
#123
Total Output Tokens
139,375
Response Time (avg)
24.87s
Total Cost
$0.478
Rank
#120
Total Output Tokens
270,259
Response Time (avg)
32.56s
Total Cost
$0.805
Recommended model KAT-Coder-Pro V2.5 (medium)

Its score stays close to the best score here (6.9 vs 7.0), while costing about 1.7x less than Grok 4.20 (medium).

Detailed comparison

Metric KAT-Coder-Pro V2.5 KAT-Coder-Pro V2.5 medium Release: 2026-07-14 Grok 4.20 Grok 4.20 medium Release: 2026-03-31
Score 6.9 7.0
Rank #123 #120
Reliability 10.0 10.0
Consistency 7.0 8.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 65.2% 60.6%
Flaky tests 8 5
Total Runs 66 66
Cost per result 4.342 10.365
Total Cost $0.478 $0.805
Input Price $0.740 / 1M $1.250 / 1M
Output Price $2.960 / 1M $2.500 / 1M
Total Input Tokens 87,916 102,986
Output Tokens 7,213 5,860
Reasoning Tokens 132,162 264,399
Response Time (avg) 24.87s 32.56s
Response Time (max) 257.00s 199.66s
Response Time (total) 547.14s 716.36s
Parameters ~700B total (~72B active) ~1T total (~100B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#123 KAT-Coder-Pro V2.5

medium
Cost
$0.009
Time
25.6s
Tokens
2,939 tok

#120 SpaceXAI: Grok 4.20

medium
Cost
$0.041
Time
110.3s
Tokens
16,336 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
KAT-Coder-Pro V2.5 7.8 10.0 66.7% 0 33.10s 7,893 416 33,658
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150

Quick Compare

Switch Comparison Pair