Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Claude Fable 5.1 (low) vs DeepSeek V3.2 (medium)

DeepSeek V3.2 (medium) unggul dalam skor rata-rata dengan 7.0 vs 6.9. DeepSeek V3.2 (medium) memiliki biaya benchmark lebih rendah di $0.078 vs $2.565. Claude Fable 5.1 (low) lebih cepat di 10.33s vs 68.43s, dengan tingkat keberhasilan 56.1% vs 65.2%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-09-02

Peringkat
#136
Total token output
24,375
Waktu respons (rata-rata)
10.33s
Total Biaya
$2.565
Peringkat
#128
Total token output
128,787
Waktu respons (rata-rata)
68.43s
Total Biaya
$0.078
Model yang direkomendasikan DeepSeek V3.2 (medium)

It has the best score here (7.0), while costing about 33.3x less than Claude Fable 5.1 (low).

Perbandingan terperinci

Metrik Claude Fable 5.1 Claude Fable 5.1 low Rilis: 2026-09-02 DeepSeek V3.2 DeepSeek V3.2 medium Rilis: 2025-12-01
Skor 6.9 7.0
Peringkat #136 #128
Keandalan 10.0 10.0
Konsistensi 9.2 7.4
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 56.1% 65.2%
Tes tidak stabil 2 7
Total Run 66 66
Biaya per hasil 23.315 0.672
Total Biaya $2.565 $0.078
Harga input $10.000 / 1M $0.269 / 1M
Harga output $50.000 / 1M $0.400 / 1M
Total token input 134,582 101,056
Token output 8,955 11,832
Token penalaran 15,420 116,955
Waktu respons (rata-rata) 10.33s 68.43s
Waktu respons (maks) 33.10s 376.10s
Waktu respons (total) 227.37s 1505.53s
Parameter ~9.5T total (~878B aktif) 671B total (37B aktif)
Ketersediaan Tertutup Sumber terbuka

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#136 Claude Fable 5.1

low
Biaya
$0.126
Waktu
28.5s
Token
2,652 tok

#128 DeepSeek V3.2

medium
Biaya
$0.001
Waktu
53.6s
Token
1,932 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Claude Fable 5.1 3.4 10.0 0.0% 0 6.75s 10,608 1,753 0
DeepSeek V3.2 6.0 7.2 55.6% 1 248.68s 5,717 649 52,014

Perbandingan Cepat

Ganti Pasangan Perbandingan