Navigasi
Advertise here

Qwen3.8 2.4T A95B (high) vs Grok 4.6 (low)

Skor rata-rata hampir imbang di 7.4 vs 7.4. Grok 4.6 (low) memiliki biaya benchmark lebih rendah di $0.449 vs $2.357. Grok 4.6 (low) lebih cepat di 14.55s vs 120.59s, dengan tingkat keberhasilan 74.2% vs 69.7%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-09-24

Model yang Dibandingkan

Peringkat
#133
Total token output
456,028
Waktu respons (rata-rata)
120.59s
Total Biaya
$2.357
Peringkat
#135
Total token output
38,048
Waktu respons (rata-rata)
14.55s
Total Biaya
$0.449
Model yang direkomendasikan Grok 4.6 (low)

It has the best score here (7.4), while costing about 5.3x less than Qwen3.8 2.4T A95B (high).

Perbandingan terperinci

Metrik Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B high Rilis: 2026-08-12 Grok 4.6 Grok 4.6 low Rilis: 2026-08-12
Skor 7.4 7.4
Peringkat #133 #135
Keandalan 9.4 10.0
Konsistensi 8.3 9.6
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 74.2% 69.7%
Tes tidak stabil 4 1
Total Run 66 66
Biaya per hasil 16.832 2.989
Total Biaya $2.357 $0.449
Harga input $2.000 / 1M $2.000 / 1M
Harga output $6.000 / 1M $6.000 / 1M
Total token input 108,931 109,960
Token output 112,801 5,124
Token penalaran 343,227 32,924
Waktu respons (rata-rata) 120.59s 14.55s
Waktu respons (maks) 532.28s 26.46s
Waktu respons (total) 2653.05s 320.17s
Parameter 2.4T total (95B aktif) ~1.7T total (~170B aktif)
Ketersediaan Bobot tersedia Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#133 Qwen3.8 2.4T A95B

high
Reached the allocated time limit (600 seconds) without receiving showcase output.
Biaya
$0.000
Waktu
600.0s
Token
0 tok

#135 SpaceXAI: Grok 4.6

low
Biaya
$0.011
Waktu
26.0s
Token
1,941 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Qwen3.8 2.4T A95B 5.9 6.3 55.6% 1 160.52s 6,399 3,310 48,078
Grok 4.6 5.5 7.1 44.4% 1 21.69s 9,579 349 7,740

Perbandingan Cepat

Ganti Pasangan Perbandingan