Navigasi
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Grok 4.5 (medium) vs GLM 5.3 (high)

Skor rata-rata hampir imbang di 7.8 vs 7.8. GLM 5.3 (high) memiliki biaya benchmark lebih rendah di $0.610 vs $2.454. GLM 5.3 (high) lebih cepat di 28.38s vs 68.44s, dengan tingkat keberhasilan 79.7% vs 75.4%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-01

Model yang Dibandingkan

Peringkat
#104
Total token output
306,044
Waktu respons (rata-rata)
68.44s
Total Biaya
$2.454
Peringkat
#105
Total token output
68,482
Waktu respons (rata-rata)
28.38s
Total Biaya
$0.610
Model yang direkomendasikan GLM 5.3 (high)

It has the best score here (7.8), while costing about 4.0x less than Grok 4.5 (medium).

Perbandingan terperinci

Metrik Grok 4.5 Grok 4.5 medium Rilis: 2026-07-08 GLM 5.3 GLM 5.3 high Rilis: 2026-08-20
Skor 7.8 7.8
Peringkat #104 #105
Keandalan 10.0 10.0
Konsistensi 9.0 7.6
Percobaan 69/69 69/69
Tes benar
Tingkat lulus per percobaan 79.7% 75.4%
Tes tidak stabil 3 7
Total Run 69 69
Biaya per hasil 14.430 4.351
Total Biaya $2.454 $0.610
Harga input $2.000 / 1M $1.400 / 1M
Harga output $6.000 / 1M $4.400 / 1M
Total token input 308,369 219,811
Token output 7,013 8,811
Token penalaran 299,031 59,671
Waktu respons (rata-rata) 68.44s 28.38s
Waktu respons (maks) 436.38s 159.91s
Waktu respons (total) 1574.14s 652.78s
Parameter ~1.7T total (~170B aktif) 744B total (40B aktif)
Ketersediaan Tertutup Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#104 SpaceXAI: Grok 4.5

medium
Biaya
$0.044
Waktu
59.4s
Token
7,512 tok

#105 GLM 5.3

high
Biaya
$0.085
Waktu
362.5s
Token
19,310 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Grok 4.5 7.6 7.2 77.8% 1 155.69s 9,579 390 104,634
GLM 5.3 6.6 5.0 66.7% 2 18.89s 7,317 473 8,190

Perbandingan Cepat

Ganti Pasangan Perbandingan