Navigasi
AI BENCHY
Advertise here

Qwen3.8 2.4T A95B (high) vs Grok Build 0.1 (medium)

Grok Build 0.1 (medium) unggul dalam skor rata-rata dengan 7.5 vs 7.4. Grok Build 0.1 (medium) memiliki biaya benchmark lebih rendah di $1.130 vs $2.357. Grok Build 0.1 (medium) lebih cepat di 54.50s vs 120.59s, dengan tingkat keberhasilan 74.2% vs 60.6%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-14

Peringkat
#92
Total token output
456,028
Waktu respons (rata-rata)
120.59s
Total Biaya
$2.357
Peringkat
#86
Total token output
511,406
Waktu respons (rata-rata)
54.50s
Total Biaya
$1.130
Model yang direkomendasikan Grok Build 0.1 (medium)

It has the best score here (7.5), while costing about 2.1x less than Qwen3.8 2.4T A95B (high).

Perbandingan terperinci

Metrik Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B high Rilis: 2026-08-12 Grok Build 0.1 Grok Build 0.1 medium Rilis: 2026-05-21
Skor 7.4 7.5
Peringkat #92 #86
Keandalan 9.4 10.0
Konsistensi 8.3 9.6
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 74.2% 60.6%
Tes tidak stabil 4 1
Total Run 66 66
Biaya per hasil 16.832 8.691
Total Biaya $2.357 $1.130
Harga input $2.000 / 1M $1.000 / 1M
Harga output $6.000 / 1M $2.000 / 1M
Total token input 108,931 106,946
Token output 112,801 7,993
Token penalaran 343,227 503,413
Waktu respons (rata-rata) 120.59s 54.50s
Waktu respons (maks) 532.28s 252.69s
Waktu respons (total) 2653.05s 1199.07s
Parameter 2.4T total (95B aktif) ~300B total (~30B aktif)
Ketersediaan Bobot tersedia Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#92 Qwen3.8 2.4T A95B

high
SVG tidak valid
Biaya
$0.000
Waktu
600.0s
Token
0 tok

#86 SpaceXAI: Grok Build 0.1

medium
Biaya
$0.028
Waktu
81.3s
Token
14,009 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Qwen3.8 2.4T A95B 5.9 6.3 55.6% 1 160.52s 6,399 3,310 48,078
Grok Build 0.1 5.7 9.7 33.3% 0 108.46s 8,304 1,138 161,452

Perbandingan Cepat

Ganti Pasangan Perbandingan