Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

LongCat 2.0 (medium) vs Grok 4.6 (low)

Skor rata-rata hampir imbang di 7.4 vs 7.4. Grok 4.6 (low) memiliki biaya benchmark lebih rendah di $0.441 vs $0.478. Grok 4.6 (low) lebih cepat di 14.20s vs 136.64s, dengan tingkat keberhasilan 60.6% vs 71.2%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-08-12

Peringkat
#88
Total token output
373,396
Waktu respons (rata-rata)
136.64s
Total Biaya
$0.478
Peringkat
#90
Total token output
36,792
Waktu respons (rata-rata)
14.20s
Total Biaya
$0.441
Model yang direkomendasikan Grok 4.6 (low)

It has the best score here (7.4), while responding about 9.6x faster than LongCat 2.0 (medium).

Perbandingan terperinci

Metrik LongCat 2.0 LongCat 2.0 medium Rilis: 2026-07-20 Grok 4.6 Grok 4.6 low Rilis: 2026-08-12
Skor 7.4 7.4
Peringkat #88 #90
Keandalan 9.8 10.0
Konsistensi 8.9 9.2
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 60.6% 71.2%
Tes tidak stabil 3 2
Total Run 66 66
Biaya per hasil 3.978 2.938
Total Biaya $0.478 $0.441
Harga input $0.300 / 1M $2.000 / 1M
Harga output $1.200 / 1M $6.000 / 1M
Total token input 97,315 109,951
Token output 44,384 5,124
Token penalaran 329,012 31,668
Waktu respons (rata-rata) 136.64s 14.20s
Waktu respons (maks) 939.52s 26.46s
Waktu respons (total) 3006.01s 312.33s
Parameter 1.6T total (48B aktif) ~1.7T total (~170B aktif)
Ketersediaan Sumber terbuka Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#88 LongCat 2.0

medium
Biaya
$0.016
Waktu
285.1s
Token
13,057 tok

#90 SpaceXAI: Grok 4.6

low
Biaya
$0.011
Waktu
26.0s
Token
1,941 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
LongCat 2.0 10.0 10.0 100.0% 0 455.00s 7,419 533 214,143
Grok 4.6 5.5 7.1 44.4% 1 21.69s 9,579 349 7,740

Perbandingan Cepat

Ganti Pasangan Perbandingan