Navigate
Advertise here

Claude Opus 4.8 vs GPT-6 Sol

Claude Opus 4.8 leads on average score with 7.3 vs 7.2. GPT-6 Sol has the lower benchmark cost at $0.280 vs $1.166. Claude Opus 4.8 is faster at 4.92s vs 6.84s, with pass rates of 63.6% vs 66.7%.

Last updated at: 2026-09-23

Compared models

Rank
#144
Total Output Tokens
16,797
Response Time (avg)
4.92s
Total Cost
$1.166
Rank
#148
Total Output Tokens
6,696
Response Time (avg)
6.84s
Total Cost
$0.280
Recommended model GPT-6 Sol

Its score stays close to the best score here (7.2 vs 7.3), while costing about 4.2x less than Claude Opus 4.8.

Detailed comparison

Metric Claude Opus 4.8 Claude Opus 4.8 none Release: 2026-05-28 GPT-6 Sol GPT-6 Sol none Release: 2026-09-23
Score 7.3 7.2
Rank #144 #148
Reliability 10.0 10.0
Consistency 9.2 8.5
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 63.6% 66.7%
Flaky tests 2 4
Total Runs 66 66
Cost per result 8.969 2.151
Total Cost $1.166 $0.280
Input Price $5.000 / 1M $2.000 / 1M
Output Price $25.000 / 1M $10.000 / 1M
Total Input Tokens 149,203 106,332
Output Tokens 16,797 6,696
Reasoning Tokens 0 0
Response Time (avg) 4.92s 6.84s
Response Time (max) 35.03s 23.89s
Response Time (total) 108.24s 150.43s
Parameters ~5T total (~500B active) ~2T total (~150B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#144 Claude Opus 4.8

none
Cost
$0.053
Time
22.0s
Tokens
2,253 tok

#148 GPT-6 Sol

none
Cost
$0.030
Time
31.9s
Tokens
3,055 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 5.5 10.0 33.3% 0 3.29s 10,590 1,332 0
GPT-6 Sol 5.5 10.0 33.3% 0 12.23s 7,302 389 0

Quick Compare

Switch Comparison Pair