Navigate
Advertise here

Mercury 2.5 (high) vs GPT-5.5

Mercury 2.5 (high) leads on average score with 7.1 vs 7.0. Mercury 2.5 (high) has the lower benchmark cost at $0.031 vs $0.544. GPT-5.5 is faster at 2.35s vs 4.32s, with pass rates of 62.1% vs 59.1%.

Last updated at: 2026-09-08

Compared models

Rank
#134
Total Output Tokens
174,547
Response Time (avg)
4.32s
Total Cost
$0.031
Rank
#139
Total Output Tokens
4,915
Response Time (avg)
2.35s
Total Cost
$0.544
Recommended model Mercury 2.5 (high)

It has the best score here (7.1), while costing about 17.6x less than GPT-5.5.

Detailed comparison

Metric Mercury 2.5 Mercury 2.5 high Release: 2026-09-08 GPT-5.5 GPT-5.5 none Release: 2026-04-24
Score 7.1 7.0
Rank #134 #139
Reliability 9.7 10.0
Consistency 8.6 9.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 62.1% 59.1%
Flaky tests 4 2
Total Runs 66 66
Cost per result 0.259 4.533
Total Cost $0.031 $0.544
Input Price $0.040 / 1M $5.000 / 1M
Output Price $0.150 / 1M $30.000 / 1M
Total Input Tokens 120,068 79,294
Output Tokens 3,402 4,915
Reasoning Tokens 171,145 0
Response Time (avg) 4.32s 2.35s
Response Time (max) 23.63s 12.24s
Response Time (total) 95.06s 51.78s
Parameters ~100B ~1.5T total (~100B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#134 Mercury 2.5

high
Cost
$0.001
Time
3.5s
Tokens
2,447 tok

#139 GPT-5.5

none
Cost
$0.090
Time
54.3s
Tokens
3,063 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 6.4 7.8 44.4% 1 6.58s 7,757 415 44,012
GPT-5.5 5.5 10.0 33.3% 0 1.35s 7,305 462 0

Quick Compare

Switch Comparison Pair