Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Grok 4.20 (medium) vs GLM 5.1 (medium)

Skor rata-rata hampir imbang di 7.0 vs 7.0. GLM 5.1 (medium) memiliki biaya benchmark lebih rendah di $0.727 vs $0.805. Grok 4.20 (medium) lebih cepat di 32.56s vs 58.80s, dengan tingkat keberhasilan 60.6% vs 65.2%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-15

Peringkat
#120
Total token output
270,259
Waktu respons (rata-rata)
32.56s
Total Biaya
$0.805
Peringkat
#117
Total token output
215,889
Waktu respons (rata-rata)
58.80s
Total Biaya
$0.727
Model yang direkomendasikan Grok 4.20 (medium)

It has the best score here (7.0), while responding about 1.8x faster than GLM 5.1 (medium).

Perbandingan terperinci

Metrik Grok 4.20 Grok 4.20 medium Rilis: 2026-03-31 GLM 5.1 GLM 5.1 medium Rilis: 2026-04-07
Skor 7.0 7.0
Peringkat #120 #117
Keandalan 10.0 8.1
Konsistensi 8.2 8.4
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 60.6% 65.2%
Tes tidak stabil 5 4
Total Run 66 66
Biaya per hasil 10.365 6.923
Total Biaya $0.805 $0.727
Harga input $1.250 / 1M $0.966 / 1M
Harga output $2.500 / 1M $3.036 / 1M
Total token input 102,986 82,592
Token output 5,860 16,089
Token penalaran 264,399 199,800
Waktu respons (rata-rata) 32.56s 58.80s
Waktu respons (maks) 199.66s 308.75s
Waktu respons (total) 716.36s 1234.72s
Parameter ~1T total (~100B aktif) 744B total (40B aktif)
Ketersediaan Tertutup Sumber terbuka

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#120 SpaceXAI: Grok 4.20

medium
Biaya
$0.041
Waktu
110.3s
Token
16,336 tok

#117 GLM 5.1

medium
SVG tidak valid
Biaya
$0.000
Waktu
300.0s
Token
0 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150
GLM 5.1 4.6 3.7 44.5% 2 109.63s 5,702 4,871 37,826

Perbandingan Cepat

Ganti Pasangan Perbandingan