Navigasi
AI BENCHY
Advertise here

DeepSeek V3.2 (medium) vs Step 3.7 Flash (low)

Step 3.7 Flash (low) unggul dalam skor rata-rata dengan 7.1 vs 7.0. DeepSeek V3.2 (medium) memiliki biaya benchmark lebih rendah di $0.078 vs $0.390. Step 3.7 Flash (low) lebih cepat di 17.87s vs 68.43s, dengan tingkat keberhasilan 65.2% vs 63.6%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-29

Peringkat
#119
Total token output
128,787
Waktu respons (rata-rata)
68.43s
Total Biaya
$0.078
Peringkat
#115
Total token output
320,393
Waktu respons (rata-rata)
17.87s
Total Biaya
$0.390
Model yang direkomendasikan DeepSeek V3.2 (medium)

Its score stays close to the best score here (7.0 vs 7.1), while costing about 5.0x less than Step 3.7 Flash (low).

Perbandingan terperinci

Metrik DeepSeek V3.2 DeepSeek V3.2 medium Rilis: 2025-12-01 Step 3.7 Flash Step 3.7 Flash low Rilis: 2026-05-29
Skor 7.0 7.1
Peringkat #119 #115
Keandalan 10.0 10.0
Konsistensi 7.4 8.1
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 65.2% 63.6%
Tes tidak stabil 7 5
Total Run 66 66
Biaya per hasil 0.672 3.539
Total Biaya $0.078 $0.390
Harga input $0.269 / 1M $0.200 / 1M
Harga output $0.400 / 1M $1.150 / 1M
Total token input 101,056 103,842
Token output 11,832 320,393
Token penalaran 116,955 0
Waktu respons (rata-rata) 68.43s 17.87s
Waktu respons (maks) 376.10s 124.75s
Waktu respons (total) 1505.53s 393.04s
Parameter 671B total (37B aktif) 198B total (11B aktif)
Ketersediaan Sumber terbuka Sumber terbuka

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#119 DeepSeek V3.2

medium
Biaya
$0.001
Waktu
53.6s
Token
1,932 tok

#115 Step 3.7 Flash

low
SVG tidak valid
Biaya
$0.004
Waktu
25.3s
Token
3,072 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
DeepSeek V3.2 6.0 7.2 55.6% 1 248.68s 5,717 649 52,014
Step 3.7 Flash 8.2 7.2 88.9% 1 9.46s 7,437 18,685 0

Perbandingan Cepat

Ganti Pasangan Perbandingan