Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

LongCat 2.0 (medium) vs GPT 5.3 Chat

GPT 5.3 Chat unggul dalam skor rata-rata dengan 7.5 vs 7.4. LongCat 2.0 (medium) memiliki biaya benchmark lebih rendah di $0.499 vs $0.533. GPT 5.3 Chat lebih cepat di 6.56s vs 143.55s, dengan tingkat keberhasilan 59.1% vs 65.2%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-09-03

Peringkat
#104
Total token output
390,873
Waktu respons (rata-rata)
143.55s
Total Biaya
$0.499
Peringkat
#93
Total token output
28,179
Waktu respons (rata-rata)
6.56s
Total Biaya
$0.533
Model yang direkomendasikan GPT 5.3 Chat

It has the best score here (7.5), while responding about 21.9x faster than LongCat 2.0 (medium).

Perbandingan terperinci

Metrik LongCat 2.0 LongCat 2.0 medium Rilis: 2026-07-20 GPT 5.3 Chat GPT 5.3 Chat none Rilis: 2026-03-03
Skor 7.4 7.5
Peringkat #104 #93
Keandalan 10.0 10.0
Konsistensi 9.3 8.6
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 59.1% 65.2%
Tes tidak stabil 2 4
Total Run 66 66
Biaya per hasil 4.153 4.097
Total Biaya $0.499 $0.533
Harga input $0.300 / 1M $1.750 / 1M
Harga output $1.200 / 1M $14.000 / 1M
Total token input 97,324 78,846
Token output 44,383 28,179
Token penalaran 346,490 0
Waktu respons (rata-rata) 143.55s 6.56s
Waktu respons (maks) 939.52s 18.33s
Waktu respons (total) 3158.11s 137.77s
Parameter 1.6T total (48B aktif) ~400B total (~17B aktif)
Ketersediaan Sumber terbuka Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#104 LongCat 2.0

medium
Biaya
$0.016
Waktu
285.1s
Token
13,057 tok

#93 GPT 5.3 Chat

none
Biaya
$0.008
Waktu
8.1s
Token
634 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
LongCat 2.0 10.0 10.0 100.0% 0 455.00s 7,419 533 214,143
GPT 5.3 Chat 5.6 4.7 55.6% 2 10.52s 7,302 6,632 0

Perbandingan Cepat

Ganti Pasangan Perbandingan