Navigasi
AI BENCHY
Advertise here

KAT-Coder-Pro V2.5 (medium) vs Grok 4.20 (medium)

Grok 4.20 (medium) unggul dalam skor rata-rata dengan 7.0 vs 6.9. KAT-Coder-Pro V2.5 (medium) memiliki biaya benchmark lebih rendah di $0.478 vs $0.805. KAT-Coder-Pro V2.5 (medium) lebih cepat di 24.87s vs 32.56s, dengan tingkat keberhasilan 65.2% vs 60.6%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-15

Peringkat
#123
Total token output
139,375
Waktu respons (rata-rata)
24.87s
Total Biaya
$0.478
Peringkat
#120
Total token output
270,259
Waktu respons (rata-rata)
32.56s
Total Biaya
$0.805
Model yang direkomendasikan KAT-Coder-Pro V2.5 (medium)

Its score stays close to the best score here (6.9 vs 7.0), while costing about 1.7x less than Grok 4.20 (medium).

Perbandingan terperinci

Metrik KAT-Coder-Pro V2.5 KAT-Coder-Pro V2.5 medium Rilis: 2026-07-14 Grok 4.20 Grok 4.20 medium Rilis: 2026-03-31
Skor 6.9 7.0
Peringkat #123 #120
Keandalan 10.0 10.0
Konsistensi 7.0 8.2
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 65.2% 60.6%
Tes tidak stabil 8 5
Total Run 66 66
Biaya per hasil 4.342 10.365
Total Biaya $0.478 $0.805
Harga input $0.740 / 1M $1.250 / 1M
Harga output $2.960 / 1M $2.500 / 1M
Total token input 87,916 102,986
Token output 7,213 5,860
Token penalaran 132,162 264,399
Waktu respons (rata-rata) 24.87s 32.56s
Waktu respons (maks) 257.00s 199.66s
Waktu respons (total) 547.14s 716.36s
Parameter ~700B total (~72B aktif) ~1T total (~100B aktif)
Ketersediaan Tertutup Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#123 KAT-Coder-Pro V2.5

medium
Biaya
$0.009
Waktu
25.6s
Token
2,939 tok

#120 SpaceXAI: Grok 4.20

medium
Biaya
$0.041
Waktu
110.3s
Token
16,336 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
KAT-Coder-Pro V2.5 7.8 10.0 66.7% 0 33.10s 7,893 416 33,658
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150

Perbandingan Cepat

Ganti Pasangan Perbandingan