Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

LongCat 2.0 (medium) vs Grok Build 0.1 (medium)

Skor rata-rata hampir imbang di 7.4 vs 7.5. LongCat 2.0 (medium) memiliki biaya benchmark lebih rendah di $0.499 vs $1.130. Grok Build 0.1 (medium) lebih cepat di 54.50s vs 143.55s, dengan tingkat keberhasilan 59.1% vs 60.6%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-15

Peringkat
#94
Total token output
390,873
Waktu respons (rata-rata)
143.55s
Total Biaya
$0.499
Peringkat
#89
Total token output
511,406
Waktu respons (rata-rata)
54.50s
Total Biaya
$1.130
Model yang direkomendasikan LongCat 2.0 (medium)

It has the best score here (7.4), while costing about 2.3x less than Grok Build 0.1 (medium).

Perbandingan terperinci

Metrik LongCat 2.0 LongCat 2.0 medium Rilis: 2026-07-20 Grok Build 0.1 Grok Build 0.1 medium Rilis: 2026-05-21
Skor 7.4 7.5
Peringkat #94 #89
Keandalan 10.0 10.0
Konsistensi 9.3 9.6
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 59.1% 60.6%
Tes tidak stabil 2 1
Total Run 66 66
Biaya per hasil 4.153 8.691
Total Biaya $0.499 $1.130
Harga input $0.300 / 1M $1.000 / 1M
Harga output $1.200 / 1M $2.000 / 1M
Total token input 97,324 106,946
Token output 44,383 7,993
Token penalaran 346,490 503,413
Waktu respons (rata-rata) 143.55s 54.50s
Waktu respons (maks) 939.52s 252.69s
Waktu respons (total) 3158.11s 1199.07s
Parameter 1.6T total (48B aktif) ~300B total (~30B aktif)
Ketersediaan Sumber terbuka Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#94 LongCat 2.0

medium
Biaya
$0.016
Waktu
285.1s
Token
13,057 tok

#89 SpaceXAI: Grok Build 0.1

medium
Biaya
$0.028
Waktu
81.3s
Token
14,009 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
LongCat 2.0 10.0 10.0 100.0% 0 455.00s 7,419 533 214,143
Grok Build 0.1 5.7 9.7 33.3% 0 108.46s 8,304 1,138 161,452

Perbandingan Cepat

Ganti Pasangan Perbandingan