Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

DeepSeek V4 Flash 0423 vs GPT-5.6 Luna

Skor rata-rata hampir imbang di 5.4 vs 5.4. GPT-5.6 Luna memiliki biaya benchmark lebih rendah di $0.029 vs $0.040. GPT-5.6 Luna lebih cepat di 1.50s vs 36.17s, dengan tingkat keberhasilan 27.3% vs 36.4%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-20

Peringkat
#218
Total token output
100,726
Waktu respons (rata-rata)
36.17s
Total Biaya
$0.040
Peringkat
#221
Total token output
6,709
Waktu respons (rata-rata)
1.50s
Total Biaya
$0.029
Model yang direkomendasikan GPT-5.6 Luna

It has the best score here (5.4), while responding about 24.1x faster than DeepSeek V4 Flash 0423.

Perbandingan terperinci

Metrik DeepSeek V4 Flash 0423 DeepSeek V4 Flash 0423 none Rilis: 2026-04-24 GPT-5.6 Luna GPT-5.6 Luna none Rilis: 2026-07-09
Skor 5.4 5.4
Peringkat #218 #221
Keandalan 10.0 10.0
Konsistensi 8.5 8.8
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 27.3% 36.4%
Tes tidak stabil 4 3
Total Run 66 66
Biaya per hasil 1.146 2.357
Total Biaya $0.040 $0.029
Harga input $0.089 / 1M $0.200 / 1M
Harga output $0.178 / 1M $1.200 / 1M
Total token input 240,230 101,332
Token output 100,726 6,709
Token penalaran 0 0
Waktu respons (rata-rata) 36.17s 1.50s
Waktu respons (maks) 247.27s 10.57s
Waktu respons (total) 795.83s 33.05s
Parameter 284B total (13B aktif) ~400B total (~17B aktif)
Ketersediaan Sumber terbuka Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#218 DeepSeek V4 Flash 0423

none
Biaya
$0.004
Waktu
157.6s
Token
11,297 tok

#221 GPT-5.6 Luna

none
Biaya
$0.016
Waktu
15.8s
Token
2,685 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
DeepSeek V4 Flash 0423 4.2 7.4 11.1% 1 17.13s 7,279 9,717 0
GPT-5.6 Luna 3.8 7.2 22.2% 1 980ms 7,302 459 0

Perbandingan Cepat

Ganti Pasangan Perbandingan