Navigasi
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Mercury 2.5 (medium) vs LongCat 2.0

Skor rata-rata hampir imbang di 6.4 vs 6.4. Mercury 2.5 (medium) memiliki biaya benchmark lebih rendah di $0.034 vs $0.104. Mercury 2.5 (medium) lebih cepat di 7.74s vs 8.50s, dengan tingkat keberhasilan 63.8% vs 42.0%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-01

Model yang Dibandingkan

Peringkat
#206
Total token output
138,946
Waktu respons (rata-rata)
7.74s
Total Biaya
$0.034
Peringkat
#211
Total token output
17,473
Waktu respons (rata-rata)
8.50s
Total Biaya
$0.104
Model yang direkomendasikan Mercury 2.5 (medium)

It has the best score here (6.4), while costing about 3.1x less than LongCat 2.0.

Perbandingan terperinci

Metrik Mercury 2.5 Mercury 2.5 medium Rilis: 2026-09-08 LongCat 2.0 LongCat 2.0 none Rilis: 2026-07-20
Skor 6.4 6.4
Peringkat #206 #211
Keandalan 9.7 10.0
Konsistensi 8.3 9.0
Percobaan 69/69 69/69
Tes benar
Tingkat lulus per percobaan 63.8% 42.0%
Tes tidak stabil 5 3
Total Run 69 69
Biaya per hasil 0.280 1.298
Total Biaya $0.034 $0.104
Harga input $0.040 / 1M $0.300 / 1M
Harga output $0.150 / 1M $1.200 / 1M
Total token input 317,982 276,199
Token output 3,918 17,473
Token penalaran 135,028 0
Waktu respons (rata-rata) 7.74s 8.50s
Waktu respons (maks) 107.89s 81.77s
Waktu respons (total) 178.07s 195.60s
Parameter ~100B 1.6T total (48B aktif)
Ketersediaan Tertutup Sumber terbuka

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#206 Mercury 2.5

medium
Biaya
$0.001
Waktu
3.8s
Token
3,386 tok

#211 LongCat 2.0

none
Biaya
$0.027
Waktu
411.7s
Token
22,283 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Mercury 2.5 5.0 5.1 44.5% 2 3.70s 7,807 488 21,857
LongCat 2.0 5.5 10.0 33.3% 0 2.85s 7,446 578 0

Perbandingan Cepat

Ganti Pasangan Perbandingan