Navigate
Advertise here

Compared models

Claude Opus 4.6 (medium) vs Claude Sonnet 4.6 (medium) vs GPT-5.3-Codex (medium) vs Gemini 3.1 Pro Preview (medium) benchmark comparison: GPT-5.3-Codex (medium) leads on Score with 9.1. Claude Opus 4.6 (medium) leads on Reliability with 10.0. GPT-5.3-Codex (medium) has the lowest Total Cost at $1.318. GPT-5.3-Codex (medium) is fastest at 18.87s.

Last updated at: 2026-10-06

Compared models

Rank
#228
Total Output Tokens
96,722
Response Time (avg)
32.23s
Total Cost
$2.961
Rank
#212
Total Output Tokens
120,427
Response Time (avg)
28.08s
Total Cost
$2.126
Rank
#34
Total Output Tokens
66,287
Response Time (avg)
18.87s
Total Cost
$1.318
Rank
#71
Total Output Tokens
102,691
Response Time (avg)
22.71s
Total Cost
$1.585
Recommended model GPT-5.3-Codex (medium)

It has the best score here (9.1), while costing about 1.7x less than the other models in this comparison.

Detailed comparison

Metric Claude Opus 4.6 Claude Opus 4.6 medium Release: 2026-02-05 Claude Sonnet 4.6 Claude Sonnet 4.6 medium Release: 2026-02-17 GPT-5.3-Codex GPT-5.3-Codex medium Release: 2026-02-05 Gemini 3.1 Pro Preview Gemini 3.1 Pro Preview medium Release: 2026-02-19
Score 6.3 6.4 9.1 8.4
Rank #228 #212 #34 #71
Reliability 10.0 10.0 10.0 9.8
Consistency 8.5 9.1 8.6 9.6
Attempts 66/69 66/69 69/69 69/69
Tests Correct
Attempt pass rate 60.9% 62.3% 84.1% 89.9%
Flaky tests 3 1 4 1
Total Runs 66 66 69 69
Cost per result 22.777 15.182 7.753 7.924
Total Cost $2.961 $2.126 $1.318 $1.585
Input Price $5.000 / 1M $3.000 / 1M $1.750 / 1M $2.000 / 1M
Output Price $25.000 / 1M $15.000 / 1M $14.000 / 1M $12.000 / 1M
Cache Read Price $0.500 / 1M $0.300 / 1M $0.175 / 1M $0.200 / 1M
Cache Write Price $6.250 / 1M $3.750 / 1M N/A $0.375 / 1M
Total Input Tokens 108,573 106,316 222,794 176,208
Output Tokens 69,994 79,338 7,357 6,093
Reasoning Tokens 26,728 41,089 58,930 96,598
Response Time (avg) 32.23s 28.08s 18.87s 22.71s
Response Time (max) 151.51s 140.96s 100.93s 88.68s
Response Time (total) 515.65s 421.15s 433.94s 386.11s
Parameters ~5T total (~500B active) ~1T total (~100B active) ~1.2T total (~80B active) ~1.2T total (~20B active)
Availability Closed Closed Closed Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#228 Claude Opus 4.6

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

#212 Claude Sonnet 4.6

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

#34 GPT-5.3-Codex

medium
Cost
$0.049
Time
54.9s
Tokens
3,580 tok

#71 Gemini 3.1 Pro Preview

medium
Cost
$0.115
Time
87.2s
Tokens
9,629 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 5.7 7.1 44.4% 1 30.10s 8,522 13,057 4,121
Claude Sonnet 4.6 5.7 6.6 44.4% 1 33.29s 6,995 16,089 3,686
GPT-5.3-Codex 10.0 10.0 100.0% 0 19.50s 7,302 535 10,890
Gemini 3.1 Pro Preview 7.9 9.9 66.7% 0 40.17s 8,124 435 41,247

Quick Compare

Switch Comparison Pair