Navigasi
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Trinity Large Thinking (high) vs GPT-5.4 Mini

Skor rata-rata hampir imbang di 5.7 vs 5.7. GPT-5.4 Mini memiliki biaya benchmark lebih rendah di $0.170 vs $0.645. GPT-5.4 Mini lebih cepat di 3.34s vs 74.52s, dengan tingkat keberhasilan 42.0% vs 30.4%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-01

Model yang Dibandingkan

Peringkat
#253
Total token output
943,044
Waktu respons (rata-rata)
74.52s
Total Biaya
$0.645
Peringkat
#255
Total token output
8,828
Waktu respons (rata-rata)
3.34s
Total Biaya
$0.170
Model yang direkomendasikan GPT-5.4 Mini

It has the best score here (5.7), while costing about 3.8x less than Trinity Large Thinking (high).

Perbandingan terperinci

Metrik Trinity Large Thinking Trinity Large Thinking high Rilis: 2026-07-28 GPT-5.4 Mini GPT-5.4 Mini none Rilis: 2026-03-17
Skor 5.7 5.7
Peringkat #253 #255
Keandalan 10.0 10.0
Konsistensi 7.3 9.3
Percobaan 69/69 69/69
Tes benar
Tingkat lulus per percobaan 42.0% 30.4%
Tes tidak stabil 8 2
Total Run 69 69
Biaya per hasil 11.289 2.818
Total Biaya $0.645 $0.170
Harga input $0.250 / 1M $0.750 / 1M
Harga output $0.800 / 1M $4.500 / 1M
Total token input 283,116 172,438
Token output 277,920 8,828
Token penalaran 665,124 0
Waktu respons (rata-rata) 74.52s 3.34s
Waktu respons (maks) 510.21s 41.98s
Waktu respons (total) 1713.90s 76.84s
Parameter 398B total (13B aktif) ~400B total (~17B aktif)
Ketersediaan Bobot tersedia Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#253 Trinity Large Thinking

high
Biaya
$0.028
Waktu
130.1s
Token
34,387 tok

#255 GPT-5.4 Mini

none
Biaya
$0.010
Waktu
11.7s
Token
2,151 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Trinity Large Thinking 3.7 4.7 33.3% 2 245.04s 7,204 83,616 266,836
GPT-5.4 Mini 5.5 10.0 33.3% 0 913ms 7,305 401 0

Perbandingan Cepat

Ganti Pasangan Perbandingan