Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

KAT-Coder-Pro V2.5 (low) vs GPT-5.6 Luna (medium)

GPT-5.6 Luna (medium) leads on average score with 7.4 vs 7.3. GPT-5.6 Luna (medium) has the lower benchmark cost at $0.072 vs $0.400. GPT-5.6 Luna (medium) is faster at 7.43s vs 20.20s, with pass rates of 68.2% vs 62.1%.

Last updated at: 2026-08-20

Rank
#100
Total Output Tokens
113,193
Response Time (avg)
20.20s
Total Cost
$0.400
Rank
#93
Total Output Tokens
44,525
Response Time (avg)
7.43s
Total Cost
$0.072
Recommended model GPT-5.6 Luna (medium)

It has the best score here (7.4), while costing about 5.6x less than KAT-Coder-Pro V2.5 (low).

Detailed comparison

Metric KAT-Coder-Pro V2.5 KAT-Coder-Pro V2.5 low Release: 2026-07-14 GPT-5.6 Luna GPT-5.6 Luna medium Release: 2026-07-09
Score 7.3 7.4
Rank #100 #93
Reliability 10.0 10.0
Consistency 7.0 9.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 68.2% 62.1%
Flaky tests 8 2
Total Runs 66 66
Cost per result 3.636 2.601
Total Cost $0.400 $0.072
Input Price $0.740 / 1M $0.200 / 1M
Output Price $2.960 / 1M $1.200 / 1M
Total Input Tokens 87,682 89,685
Output Tokens 7,166 5,701
Reasoning Tokens 106,027 38,824
Response Time (avg) 20.20s 7.43s
Response Time (max) 209.15s 29.85s
Response Time (total) 444.35s 163.54s
Parameters ~700B total (~72B active) ~400B total (~17B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#100 KAT-Coder-Pro V2.5

low
Cost
$0.016
Time
47.6s
Tokens
5,182 tok

#93 GPT-5.6 Luna

medium
Cost
$0.014
Time
14.6s
Tokens
2,362 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
KAT-Coder-Pro V2.5 7.8 10.0 66.7% 0 24.87s 7,893 402 24,945
GPT-5.6 Luna 5.4 7.2 44.4% 1 10.38s 7,302 564 9,928

Quick Compare

Switch Comparison Pair