Navigasi
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Mercury 2.5 Preview (medium) vs GPT-6 Sol

Skor rata-rata hampir imbang di 6.8 vs 6.8. Mercury 2.5 Preview (medium) memiliki biaya benchmark lebih rendah di $0.037 vs $0.551. Mercury 2.5 Preview (medium) lebih cepat di 6.66s vs 9.02s, dengan tingkat keberhasilan 66.7% vs 63.8%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-01

Model yang Dibandingkan

Peringkat
#171
Total token output
150,107
Waktu respons (rata-rata)
6.66s
Total Biaya
$0.037
Peringkat
#172
Total token output
7,989
Waktu respons (rata-rata)
9.02s
Total Biaya
$0.551
Model yang direkomendasikan Mercury 2.5 Preview (medium)

It has the best score here (6.8), while costing about 15.0x less than GPT-6 Sol.

Perbandingan terperinci

Metrik Mercury 2.5 Preview Mercury 2.5 Preview medium Rilis: 2026-09-02 GPT-6 Sol GPT-6 Sol none Rilis: 2026-09-23
Skor 6.8 6.8
Peringkat #171 #172
Keandalan 10.0 10.0
Konsistensi 9.6 8.5
Percobaan 69/69 69/69
Tes benar
Tingkat lulus per percobaan 66.7% 63.8%
Tes tidak stabil 1 4
Total Run 69 69
Biaya per hasil 0.244 4.232
Total Biaya $0.037 $0.551
Harga input $0.000 / 1M $2.000 / 1M
Harga output $0.000 / 1M $10.000 / 1M
Total token input 397,291 235,086
Token output 6,355 7,989
Token penalaran 143,752 0
Waktu respons (rata-rata) 6.66s 9.02s
Waktu respons (maks) 74.63s 57.02s
Waktu respons (total) 153.29s 207.45s
Parameter ~100B ~2T total (~150B aktif)
Ketersediaan Tertutup Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#171 Mercury 2.5 Preview

medium
Biaya
$0.001
Waktu
2.7s
Token
2,241 tok

#172 GPT-6 Sol

none
Biaya
$0.030
Waktu
31.9s
Token
3,055 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Mercury 2.5 Preview 7.8 10.0 66.7% 0 3.86s 7,839 552 22,126
GPT-6 Sol 5.5 10.0 33.3% 0 12.23s 7,302 389 0

Perbandingan Cepat

Ganti Pasangan Perbandingan