Navigasi
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Mercury 2.5 Preview (medium) vs Qwen3.7 Max

Skor rata-rata hampir imbang di 7.5 vs 7.4. Mercury 2.5 Preview (medium) memiliki biaya benchmark lebih rendah di $0.024 vs $0.197. Mercury 2.5 Preview (medium) lebih cepat di 3.58s vs 4.56s, dengan tingkat keberhasilan 69.7% vs 68.2%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-09-02

Peringkat
#91
Total token output
127,034
Waktu respons (rata-rata)
3.58s
Total Biaya
$0.024
Peringkat
#95
Total token output
12,446
Waktu respons (rata-rata)
4.56s
Total Biaya
$0.197
Model yang direkomendasikan Mercury 2.5 Preview (medium)

It has the best score here (7.5), while costing about 8.3x less than Qwen3.7 Max.

Perbandingan terperinci

Metrik Mercury 2.5 Preview Mercury 2.5 Preview medium Rilis: 2026-09-02 Qwen3.7 Max Qwen3.7 Max none Rilis: 2026-05-22
Skor 7.5 7.4
Peringkat #91 #95
Keandalan 10.0 9.9
Konsistensi 9.6 10.0
Percobaan 66/66 66/66
Tes benar
Tingkat lulus per percobaan 69.7% 68.2%
Tes tidak stabil 1 0
Total Run 66 66
Biaya per hasil 0.158 1.581
Total Biaya $0.024 $0.197
Harga input $0.040 / 1M $1.475 / 1M
Harga output $0.150 / 1M $4.425 / 1M
Total token input 112,713 95,992
Token output 5,436 12,446
Token penalaran 121,598 0
Waktu respons (rata-rata) 3.58s 4.56s
Waktu respons (maks) 17.84s 72.30s
Waktu respons (total) 78.66s 100.30s
Parameter ~100B ~1T total (~40B aktif)
Ketersediaan Tertutup Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#91 Mercury 2.5 Preview

medium
Biaya
$0.001
Waktu
2.7s
Token
2,241 tok

#95 Qwen3.7 Max

none
Biaya
$0.046
Waktu
195.0s
Token
12,171 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Mercury 2.5 Preview 7.8 10.0 66.7% 0 3.86s 7,839 552 22,126
Qwen3.7 Max 5.5 10.0 33.3% 0 1.35s 7,911 582 0

Perbandingan Cepat

Ganti Pasangan Perbandingan