Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Seed-2.0-Lite (medium) vs DeepSeek V4 Flash 0731 (high)

DeepSeek V4 Flash 0731 (high) unggul dalam skor rata-rata dengan 8.0 vs 7.9. DeepSeek V4 Flash 0731 (high) memiliki biaya benchmark lebih rendah di $0.123 vs $0.234. Seed-2.0-Lite (medium) lebih cepat di 48.53s vs 110.77s, dengan tingkat keberhasilan 74.2% vs 78.8%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-01

Peringkat
#44
Total token output
100,580
Waktu respons (rata-rata)
48.53s
Total Biaya
$0.234
Peringkat
#40
Total token output
544,164
Waktu respons (rata-rata)
110.77s
Total Biaya
$0.123
Model yang direkomendasikan DeepSeek V4 Flash 0731 (high)

It has the best score here (8.0), while costing about 1.9x less than Seed-2.0-Lite (medium).

Perbandingan terperinci

Metrik Seed-2.0-Lite Seed-2.0-Lite medium Rilis: 2026-02-14 DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 high Rilis: 2026-08-01
Skor 7.9 8.0
Peringkat #44 #40
Keandalan 10.0 9.6
Konsistensi 8.6 7.8
Tes benar
Tingkat lulus per percobaan 74.2% 78.8%
Tes tidak stabil 4 6
Total Run 66 66
Biaya per hasil 1.669 0.878
Total Biaya $0.234 $0.123
Harga input $0.250 / 1M $0.140 / 1M
Harga output $2.000 / 1M $0.280 / 1M
Total token input 129,897 96,609
Token output 12,533 32,924
Token penalaran 88,047 511,240
Waktu respons (rata-rata) 48.53s 110.77s
Waktu respons (maks) 254.92s 600.61s
Waktu respons (total) 1067.74s 2437.00s

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#44 Seed-2.0-Lite

medium
Biaya
$0.005
Waktu
86.7s
Token
2,354 tok

#40 DeepSeek V4 Flash 0731

high
SVG tidak valid
Biaya
$0.000
Waktu
187.3s
Token
19,048 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Seed-2.0-Lite 8.0 9.8 66.7% 0 156.74s 8,247 458 31,890
DeepSeek V4 Flash 0731 6.3 4.2 77.8% 2 252.67s 7,044 2,160 191,210

Perbandingan Cepat

Ganti Pasangan Perbandingan