Navigasi
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

MiniMax M2.5 (medium) vs gpt-oss-120b

MiniMax M2.5 (medium) unggul dalam skor rata-rata dengan 4.3 vs 4.2. gpt-oss-120b memiliki biaya benchmark lebih rendah di $0.016 vs $0.440. gpt-oss-120b lebih cepat di 28.63s vs 102.32s, dengan tingkat keberhasilan 42.0% vs 34.8%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-07

Model yang Dibandingkan

Peringkat
#359
Total token output
284,991
Waktu respons (rata-rata)
102.32s
Total Biaya
$0.440
Peringkat
#361
Total token output
62,213
Waktu respons (rata-rata)
28.63s
Total Biaya
$0.016
Model yang direkomendasikan gpt-oss-120b

Its score stays close to the best score here (4.2 vs 4.3), while costing about 28.4x less than MiniMax M2.5 (medium).

Perbandingan terperinci

Metrik MiniMax M2.5 MiniMax M2.5 medium Rilis: 2026-02-12 gpt-oss-120b gpt-oss-120b none Rilis: 2025-08-05
Skor 4.3 4.2
Peringkat #359 #361
Keandalan 9.3 10.0
Konsistensi 6.8 7.6
Percobaan 69/69 60/69
Tes benar
Tingkat lulus per percobaan 42.0% 34.8%
Tes tidak stabil 9 3
Total Run 69 60
Biaya per hasil 7.104 0.274
Total Biaya $0.440 $0.016
Harga input $0.270 / 1M $0.037 / 1M
Harga output $1.080 / 1M $0.170 / 1M
Harga baca cache $0.027 / 1M T/A
Harga tulis cache T/A T/A
Total token input 188,239 131,652
Token output 27,822 16,869
Token penalaran 262,674 53,081
Waktu respons (rata-rata) 102.32s 28.63s
Waktu respons (maks) 600.81s 140.90s
Waktu respons (total) 1637.05s 486.69s
Parameter 230B total (10B aktif) 117B total (5.1B aktif)
Ketersediaan Bobot tersedia Sumber terbuka

Harga cache berlaku untuk token input. Pembacaan memakai kembali prompt tersimpan; penulisan menyimpannya dan dapat menambah biaya. Token output menggunakan harga output.

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#359 MiniMax M2.5

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Biaya
$0.000
Waktu
300.0s
Token
0 tok

#361 gpt-oss-120b

none
Belum ada hasil showcase yang dihasilkan untuk model ini.
Biaya
T/A
Waktu
-
Token
0 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
MiniMax M2.5 3.4 9.1 0.0% 0 188.58s 6,076 357 106,177
gpt-oss-120b 1.5 0.4 22.2% 1 9.57s 901 204 3,028

Ganti Pasangan Perbandingan