Navigasi
AI BENCHY
Advertise here

Meituan: LongCat 2.0 vs Qwen: Qwen3.6 Max Preview

LongCat 2.0 (low) unggul dalam skor rata-rata dengan 6.7 vs 6.6. Qwen3.6 Max Preview memiliki biaya benchmark lebih rendah di $0.231 vs $0.391. Qwen3.6 Max Preview lebih cepat di 7.82s vs 100.31s, dengan tingkat keberhasilan 53.0% vs 60.6%.

Model yang direkomendasikanQwen3.6 Max PreviewIts score stays close to the best score here (6.6 vs 6.7), while costing about 1.7x less than LongCat 2.0 (low).

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-07-20

Metrik LongCat 2.0 LongCat 2.0 low Rilis: 2026-07-20 Qwen3.6 Max Preview Qwen3.6 Max Preview none Rilis: 2026-04-20
Skor 6.7 6.6
Peringkat #91 #98
Keandalan 9.6 9.9
Konsistensi 8.9 9.3
Tes benar
Tingkat lulus per percobaan 53.0% 60.6%
Tes tidak stabil 3 2
Total Run 66 66
Biaya per hasil 3.910 2.061
Total Biaya $0.391 $0.231
Harga input $0.300 / 1M $1.040 / 1M
Harga output $1.200 / 1M $6.240 / 1M
Total token input 90,909 106,339
Token output 5,679 19,257
Token penalaran 297,369 0
Waktu respons (rata-rata) 100.31s 7.82s
Waktu respons (maks) 560.39s 102.62s
Waktu respons (total) 2206.82s 172.01s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#91 LongCat 2.0

low
Biaya
$0.024
Waktu
428.0s
Token
19,557 tok

#98 Qwen3.6 Max Preview

none
Biaya
$0.025
Waktu
83.9s
Token
4,066 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
LongCat 2.0 6.6 4.6 77.8% 2 479.30s 6,544 461 206,627
Qwen3.6 Max Preview 3.8 7.3 22.2% 1 3.12s 7,913 456 0

Perbandingan Cepat

Ganti Pasangan Perbandingan