Navigasi
Advertise here

Model yang Dibandingkan

Perbandingan benchmark Claude Opus 4.6 (medium) vs Claude Sonnet 4.6 (medium) vs GPT-5.3-Codex (medium) vs Gemini 3.1 Pro Preview (medium): GPT-5.3-Codex (medium) unggul pada Skor dengan 9.1. Claude Opus 4.6 (medium) unggul pada Keandalan dengan 10.0. GPT-5.3-Codex (medium) memiliki Total Biaya terendah di $1.318. GPT-5.3-Codex (medium) paling cepat di 18.87s.

Benchmark dihasilkan dari suite pengujian AI BENCHY pada: 2026-10-06

Model yang Dibandingkan

Peringkat
#228
Total token output
96,722
Waktu respons (rata-rata)
32.23s
Total Biaya
$2.961
Peringkat
#212
Total token output
120,427
Waktu respons (rata-rata)
28.08s
Total Biaya
$2.126
Peringkat
#34
Total token output
66,287
Waktu respons (rata-rata)
18.87s
Total Biaya
$1.318
Peringkat
#71
Total token output
102,691
Waktu respons (rata-rata)
22.71s
Total Biaya
$1.585
Model yang direkomendasikan GPT-5.3-Codex (medium)

It has the best score here (9.1), while costing about 1.7x less than model lain dalam perbandingan ini.

Perbandingan terperinci

Metrik Claude Opus 4.6 Claude Opus 4.6 medium Rilis: 2026-02-05 Claude Sonnet 4.6 Claude Sonnet 4.6 medium Rilis: 2026-02-17 GPT-5.3-Codex GPT-5.3-Codex medium Rilis: 2026-02-05 Gemini 3.1 Pro Preview Gemini 3.1 Pro Preview medium Rilis: 2026-02-19
Skor 6.3 6.4 9.1 8.4
Peringkat #228 #212 #34 #71
Keandalan 10.0 10.0 10.0 9.8
Konsistensi 8.5 9.1 8.6 9.6
Percobaan 66/69 66/69 69/69 69/69
Tes benar
Tingkat lulus per percobaan 60.9% 62.3% 84.1% 89.9%
Tes tidak stabil 3 1 4 1
Total Run 66 66 69 69
Biaya per hasil 22.777 15.182 7.753 7.924
Total Biaya $2.961 $2.126 $1.318 $1.585
Harga input $5.000 / 1M $3.000 / 1M $1.750 / 1M $2.000 / 1M
Harga output $25.000 / 1M $15.000 / 1M $14.000 / 1M $12.000 / 1M
Harga baca cache $0.500 / 1M $0.300 / 1M $0.175 / 1M $0.200 / 1M
Harga tulis cache $6.250 / 1M $3.750 / 1M T/A $0.375 / 1M
Total token input 108,573 106,316 222,794 176,208
Token output 69,994 79,338 7,357 6,093
Token penalaran 26,728 41,089 58,930 96,598
Waktu respons (rata-rata) 32.23s 28.08s 18.87s 22.71s
Waktu respons (maks) 151.51s 140.96s 100.93s 88.68s
Waktu respons (total) 515.65s 421.15s 433.94s 386.11s
Parameter ~5T total (~500B aktif) ~1T total (~100B aktif) ~1.2T total (~80B aktif) ~1.2T total (~20B aktif)
Ketersediaan Tertutup Tertutup Tertutup Tertutup

Harga cache berlaku untuk token input. Pembacaan memakai kembali prompt tersimpan; penulisan menyimpannya dan dapat menambah biaya. Token output menggunakan harga output.

Showcase generasi model

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#228 Claude Opus 4.6

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Biaya
$0.000
Waktu
300.0s
Token
0 tok

#212 Claude Sonnet 4.6

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Biaya
$0.000
Waktu
300.0s
Token
0 tok

#34 GPT-5.3-Codex

medium
Biaya
$0.049
Waktu
54.9s
Token
3,580 tok

#71 Gemini 3.1 Pro Preview

medium
Biaya
$0.115
Waktu
87.2s
Token
9,629 tok

Model teratas berdasarkan skor

Skor vs Total Biaya

Waktu respons (rata-rata)

Skor vs Waktu respons (rata-rata)

Total token output

Skor vs Total token output

Rincian Kategori

Pemrograman Skor Konsistensi Tingkat lulus per percobaan Tes tidak stabil Tes benar Waktu respons (rata-rata) Token input Token output Token penalaran
Claude Opus 4.6 5.7 7.1 44.4% 1 30.10s 8,522 13,057 4,121
Claude Sonnet 4.6 5.7 6.6 44.4% 1 33.29s 6,995 16,089 3,686
GPT-5.3-Codex 10.0 10.0 100.0% 0 19.50s 7,302 535 10,890
Gemini 3.1 Pro Preview 7.9 9.9 66.7% 0 40.17s 8,124 435 41,247

Perbandingan Cepat

Ganti Pasangan Perbandingan