Navigasi
Advertise here

Trinity Large Thinking (low) vs Space Bunny Alpha (medium)

Skor rata-rata hampir imbang di 6.3 vs 6.3. Space Bunny Alpha (medium) memiliki biaya benchmark lebih rendah di $0.000 vs $0.700. Space Bunny Alpha (medium) lebih cepat di 32.88s vs 95.16s, dengan tingkat keberhasilan 47.8% vs 55.1%.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-01

Model yang Dibandingkan

Peringkat
#222
Total token output
842,200
Waktu respons (rata-rata)
95.16s
Total Biaya
$0.700
Peringkat
#220
Total token output
192,115
Waktu respons (rata-rata)
32.88s
Total Biaya
$0.000
Model yang direkomendasikan Trinity Large Thinking (low)

It has the strongest score in this comparison (6.3) and the best overall balance of cost and response time across all 2 models.

Perbandingan terperinci

Metrik Trinity Large Thinking Trinity Large Thinking low Rilis: 2026-07-28 Space Bunny Alpha Space Bunny Alpha medium Rilis: 2026-09-24
Skor 6.3 6.3
Peringkat #222 #220
Keandalan 10.0 10.0
Konsistensi 8.1 8.2
Percobaan 69/69 69/69
Tes benar
Tingkat lulus per percobaan 47.8% 55.1%
Tes tidak stabil 6 5
Total Run 69 69
Biaya per hasil 9.163 0.000
Total Biaya $0.700 $0.000
Harga input $0.250 / 1M $0.000 / 1M
Harga output $0.800 / 1M $0.000 / 1M
Total token input 347,980 267,840
Token output 121,898 192,115
Token penalaran 720,302 0
Waktu respons (rata-rata) 95.16s 32.88s
Waktu respons (maks) 540.96s 197.32s
Waktu respons (total) 2188.74s 756.28s
Parameter 398B total (13B aktif) -
Ketersediaan Bobot tersedia Tertutup

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#222 Trinity Large Thinking

low
Biaya
$0.021
Waktu
173.2s
Token
24,586 tok

#220 Space Bunny Alpha

medium
Biaya
$0.000
Waktu
35.4s
Token
4,894 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Trinity Large Thinking 5.3 10.0 33.3% 0 344.45s 5,441 27,039 217,241
Space Bunny Alpha 5.4 7.2 44.4% 1 6.82s 8,493 6,691 0

Perbandingan Cepat

Ganti Pasangan Perbandingan