Navigasi
Advertise here

Qwen3.8 2.4T A95B (low) vs Grok 4.7 (medium)

Skor rata-rata hampir imbang di 8.1 vs 8.1. Grok 4.7 (medium) memiliki biaya benchmark lebih rendah di $2.254 vs $2.672. Grok 4.7 (medium) lebih cepat di 106.09s vs 118.97s, dengan tingkat keberhasilan 72.7% vs 81.8%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-09-21

Model yang Dibandingkan

Peringkat
#70
Total token output
495,541
Waktu respons (rata-rata)
118.97s
Total Biaya
$2.672
Peringkat
#73
Total token output
433,733
Waktu respons (rata-rata)
106.09s
Total Biaya
$2.254
Model yang direkomendasikan Grok 4.7 (medium)

It has the strongest score in this comparison (8.1) and the best overall balance of cost and response time across all 2 models.

Perbandingan terperinci

Metrik Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B low Rilis: 2026-08-12 Grok 4.7 Grok 4.7 medium Rilis: 2026-09-21
Skor 8.1 8.1
Peringkat #70 #73
Keandalan 9.4 9.9
Konsistensi 9.2 8.3
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 72.7% 81.8%
Tes tidak stabil 2 5
Total Run 66 66
Biaya per hasil 17.812 15.023
Total Biaya $2.672 $2.254
Harga input $2.000 / 1M $1.600 / 1M
Harga output $6.000 / 1M $4.800 / 1M
Total token input 113,279 107,193
Token output 121,045 4,661
Token penalaran 374,496 429,072
Waktu respons (rata-rata) 118.97s 106.09s
Waktu respons (maks) 534.21s 660.89s
Waktu respons (total) 2617.41s 2333.94s
Parameter 2.4T total (95B aktif) ~1.7T total (~170B aktif)
Ketersediaan Bobot tersedia Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#70 Qwen3.8 2.4T A95B

low
Biaya
$0.100
Waktu
193.2s
Token
16,767 tok

#73 SpaceXAI: Grok 4.7

medium
Biaya
$0.047
Waktu
132.3s
Token
9,954 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Qwen3.8 2.4T A95B 7.8 9.3 66.7% 0 172.53s 6,716 21,405 57,399
Grok 4.7 6.6 4.6 77.8% 2 312.17s 10,266 349 172,347

Perbandingan Cepat

Ganti Pasangan Perbandingan