Navigate
Advertise here

Mercury 2.5 (high) vs GPT-5.6 Sol

The average score is effectively tied at 7.1 vs 7.0. Mercury 2.5 (high) has the lower benchmark cost at $0.031 vs $0.201. GPT-5.6 Sol is faster at 2.17s vs 4.32s, with pass rates of 62.1% vs 63.6%.

Last updated at: 2026-09-08

Compared models

Rank
#134
Total Output Tokens
174,547
Response Time (avg)
4.32s
Total Cost
$0.031
Rank
#136
Total Output Tokens
4,357
Response Time (avg)
2.17s
Total Cost
$0.201
Recommended model Mercury 2.5 (high)

It has the best score here (7.1), while costing about 6.5x less than GPT-5.6 Sol.

Detailed comparison

Metric Mercury 2.5 Mercury 2.5 high Release: 2026-09-08 GPT-5.6 Sol GPT-5.6 Sol none Release: 2026-07-09
Score 7.1 7.0
Rank #134 #136
Reliability 9.7 10.0
Consistency 8.6 9.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 62.1% 63.6%
Flaky tests 4 3
Total Runs 66 66
Cost per result 0.259 4.365
Total Cost $0.031 $0.201
Input Price $0.040 / 1M $2.000 / 1M
Output Price $0.150 / 1M $10.000 / 1M
Total Input Tokens 120,068 78,602
Output Tokens 3,402 4,357
Reasoning Tokens 171,145 0
Response Time (avg) 4.32s 2.17s
Response Time (max) 23.63s 12.81s
Response Time (total) 95.06s 47.66s
Parameters ~100B ~2T total (~150B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#134 Mercury 2.5

high
Cost
$0.001
Time
3.5s
Tokens
2,447 tok

#136 GPT-5.6 Sol

none
Cost
$0.116
Time
78.7s
Tokens
3,944 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2.5 6.4 7.8 44.4% 1 6.58s 7,757 415 44,012
GPT-5.6 Sol 5.5 10.0 33.3% 0 1.39s 7,302 390 0

Quick Compare

Switch Comparison Pair