Navigasi
AI BENCHY
Advertise here

DeepSeek V3.2 (medium) vs Grok 4.20 (medium)

Skor rata-rata hampir imbang di 7.0 vs 7.0. DeepSeek V3.2 (medium) memiliki biaya benchmark lebih rendah di $0.078 vs $0.805. Grok 4.20 (medium) lebih cepat di 32.56s vs 68.43s, dengan tingkat keberhasilan 65.2% vs 60.6%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-15

Peringkat
#116
Total token output
128,787
Waktu respons (rata-rata)
68.43s
Total Biaya
$0.078
Peringkat
#120
Total token output
270,259
Waktu respons (rata-rata)
32.56s
Total Biaya
$0.805
Model yang direkomendasikan DeepSeek V3.2 (medium)

It has the best score here (7.0), while costing about 10.4x less than Grok 4.20 (medium).

Perbandingan terperinci

Metrik DeepSeek V3.2 DeepSeek V3.2 medium Rilis: 2025-12-01 Grok 4.20 Grok 4.20 medium Rilis: 2026-03-31
Skor 7.0 7.0
Peringkat #116 #120
Keandalan 10.0 10.0
Konsistensi 7.4 8.2
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 65.2% 60.6%
Tes tidak stabil 7 5
Total Run 66 66
Biaya per hasil 0.672 10.365
Total Biaya $0.078 $0.805
Harga input $0.269 / 1M $1.250 / 1M
Harga output $0.400 / 1M $2.500 / 1M
Total token input 101,056 102,986
Token output 11,832 5,860
Token penalaran 116,955 264,399
Waktu respons (rata-rata) 68.43s 32.56s
Waktu respons (maks) 376.10s 199.66s
Waktu respons (total) 1505.53s 716.36s
Parameter 671B total (37B aktif) ~1T total (~100B aktif)
Ketersediaan Sumber terbuka Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#116 DeepSeek V3.2

medium
Biaya
$0.001
Waktu
53.6s
Token
1,932 tok

#120 SpaceXAI: Grok 4.20

medium
Biaya
$0.041
Waktu
110.3s
Token
16,336 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
DeepSeek V3.2 6.0 7.2 55.6% 1 248.68s 5,717 649 52,014
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150

Perbandingan Cepat

Ganti Pasangan Perbandingan