Navigate
Advertise here

Compared models

GPT-5.5 (medium) vs GPT-5.4 (medium) vs Gemini 3.1 Pro Preview (medium) vs Claude Opus 4.7 (medium) benchmark comparison: Gemini 3.1 Pro Preview (medium) leads on Score with 9.2. GPT-5.5 (medium) leads on Reliability with 10.0. Gemini 3.1 Pro Preview (medium) has the lowest Total Cost at $1.352. Claude Opus 4.7 (medium) is fastest at 7.62s.

Last updated at: 2026-09-28

Compared models

Rank
#33
Total Output Tokens
125,316
Response Time (avg)
38.97s
Total Cost
$4.163
Rank
#56
Total Output Tokens
90,466
Response Time (avg)
22.85s
Total Cost
$1.560
Rank
#22
Total Output Tokens
97,238
Response Time (avg)
20.62s
Total Cost
$1.352
Rank
#49
Total Output Tokens
29,990
Response Time (avg)
7.62s
Total Cost
$1.476
Recommended model Gemini 3.1 Pro Preview (medium)

It has the best score here (9.2), while costing about 1.8x less than the other models in this comparison.

Detailed comparison

Metric GPT-5.5 GPT-5.5 medium Release: 2026-04-24 GPT-5.4 GPT-5.4 medium Release: 2026-03-05 Gemini 3.1 Pro Preview Gemini 3.1 Pro Preview medium Release: 2026-02-19 Claude Opus 4.7 Claude Opus 4.7 medium Release: 2026-04-16
Score 9.0 8.5 9.2 8.7
Rank #33 #56 #22 #49
Reliability 10.0 10.0 10.0 10.0
Consistency 8.9 8.6 10.0 9.6
Attempts 66/66 66/66 66/66 66/66
Tests Correct
Attempt pass rate 87.9% 77.3% 90.9% 83.3%
Flaky tests 3 4 0 1
Total Runs 66 66 66 66
Cost per result 23.127 10.399 6.758 8.200
Total Cost $4.163 $1.560 $1.352 $1.476
Input Price $5.000 / 1M $2.500 / 1M $2.000 / 1M $5.000 / 1M
Output Price $30.000 / 1M $15.000 / 1M $12.000 / 1M $25.000 / 1M
Total Input Tokens 80,668 81,136 92,296 145,249
Output Tokens 5,617 6,155 5,232 24,948
Reasoning Tokens 119,699 84,311 92,006 5,042
Response Time (avg) 38.97s 22.85s 20.62s 7.62s
Response Time (max) 332.10s 100.41s 88.68s 65.40s
Response Time (total) 857.43s 502.78s 329.94s 159.94s
Parameters ~1.5T total (~100B active) ~1.5T total (~100B active) ~1.2T total (~20B active) ~5T total (~500B active)
Availability Closed Closed Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#33 GPT-5.5

medium
Cost
$0.112
Time
71.9s
Tokens
3,807 tok

#56 GPT-5.4

medium
Cost
$0.214
Time
199.6s
Tokens
14,349 tok

#22 Gemini 3.1 Pro Preview

medium
Cost
$0.115
Time
87.2s
Tokens
9,629 tok

#49 Claude Opus 4.7

medium
Cost
$0.059
Time
26.8s
Tokens
2,475 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.5 8.8 7.8 88.9% 1 59.77s 7,305 362 24,959
GPT-5.4 8.8 7.8 88.9% 1 44.36s 7,305 433 24,216
Gemini 3.1 Pro Preview 7.9 9.9 66.7% 0 40.17s 8,124 435 41,247
Claude Opus 4.7 7.6 7.2 77.8% 1 12.96s 10,635 7,629 1,114

Quick Compare

Switch Comparison Pair