Navigate
Advertise here

GPT-5.6 Sol vs Grok 4.7 (xhigh)

Grok 4.7 (xhigh) leads on average score with 6.7 vs 6.7. GPT-5.6 Sol has the lower benchmark cost at $0.472 vs $4.054. GPT-5.6 Sol is faster at 4.61s vs 123.96s, with pass rates of 60.9% vs 73.9%.

Last updated at: 2026-10-01

Compared models

Rank
#183
Total Output Tokens
5,555
Response Time (avg)
4.61s
Total Cost
$0.472
Rank
#175
Total Output Tokens
572,315
Response Time (avg)
123.96s
Total Cost
$4.054
Recommended model GPT-5.6 Sol

Its score stays close to the best score here (6.7 vs 6.7), while costing about 8.6x less than Grok 4.7 (xhigh).

Detailed comparison

Metric GPT-5.6 Sol GPT-5.6 Sol none Release: 2026-07-09 Grok 4.7 Grok 4.7 xhigh Release: 2026-09-21
Score 6.7 6.7
Rank #183 #175
Reliability 10.0 9.7
Consistency 9.1 7.8
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 60.9% 73.9%
Flaky tests 3 6
Total Runs 69 69
Cost per result 6.619 24.272
Total Cost $0.472 $4.054
Input Price $2.000 / 1M $2.000 / 1M
Output Price $10.000 / 1M $6.000 / 1M
Total Input Tokens 207,852 310,049
Output Tokens 5,555 6,094
Reasoning Tokens 0 566,221
Response Time (avg) 4.61s 123.96s
Response Time (max) 58.28s 775.64s
Response Time (total) 105.93s 2851.07s
Parameters ~2T total (~150B active) ~1.7T total (~170B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#183 GPT-5.6 Sol

none
Cost
$0.116
Time
78.7s
Tokens
3,944 tok

#175 SpaceXAI: Grok 4.7

xhigh
Cost
$0.235
Time
544.2s
Tokens
39,247 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Sol 5.5 10.0 33.3% 0 1.39s 7,302 390 0
Grok 4.7 5.3 4.5 55.6% 2 414.53s 10,168 327 256,935

Quick Compare

Switch Comparison Pair