Navigasi
Advertise here

Trinity Large Thinking (high) vs Mistral Large 4

Skor rata-rata hampir imbang di 5.7 vs 5.7. Mistral Large 4 memiliki biaya benchmark lebih rendah di $0.290 vs $0.794. Mistral Large 4 lebih cepat di 20.43s vs 74.52s, dengan tingkat keberhasilan 42.0% vs 31.9%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-07

Model yang Dibandingkan

Peringkat
#269
Total token output
717,762
Waktu respons (rata-rata)
74.52s
Total Biaya
$0.794
Peringkat
#272
Total token output
54,388
Waktu respons (rata-rata)
20.43s
Total Biaya
$0.290
Model yang direkomendasikan Mistral Large 4

It has the best score here (5.7), while costing about 2.7x less than Trinity Large Thinking (high).

Perbandingan terperinci

Metrik Trinity Large Thinking Trinity Large Thinking high Rilis: 2026-07-28 Mistral Large 4 Mistral Large 4 none Rilis: 2026-10-06
Skor 5.7 5.7
Peringkat #269 #272
Keandalan 10.0 9.7
Konsistensi 7.3 9.3
Percobaan 69/69 69/69
Tes benar
Tingkat lulus per percobaan 42.0% 31.9%
Tes tidak stabil 8 2
Total Run 69 69
Biaya per hasil 11.289 4.818
Total Biaya $0.794 $0.290
Harga input $0.250 / 1M $0.680 / 1M
Harga output $0.800 / 1M $2.090 / 1M
Harga baca cache $0.060 / 1M $0.070 / 1M
Harga tulis cache T/A T/A
Total token input 283,116 257,872
Token output 277,920 54,388
Token penalaran 665,124 0
Waktu respons (rata-rata) 74.52s 20.43s
Waktu respons (maks) 510.21s 192.80s
Waktu respons (total) 1713.90s 469.88s
Parameter 398B total (13B aktif) 1.05T total (49B aktif)
Ketersediaan Bobot tersedia Tertutup

Harga cache berlaku untuk token input. Pembacaan memakai kembali prompt tersimpan; penulisan menyimpannya dan dapat menambah biaya. Token output menggunakan harga output.

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#269 Trinity Large Thinking

high
Biaya
$0.028
Waktu
130.1s
Token
34,387 tok

#272 Mistral Large 4

none
Biaya
$0.005
Waktu
33.0s
Token
2,461 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Trinity Large Thinking 3.7 4.7 33.3% 2 245.04s 7,204 83,616 266,836
Mistral Large 4 4.7 7.5 22.2% 1 85.04s 7,431 39,805 0

Ganti Pasangan Perbandingan