Navigate
Advertise here

GPT-6.1 Sol (xhigh) vs Grok 4.6 (high)

GPT-6.1 Sol (xhigh) leads on average score with 9.9 vs 9.3. GPT-6.1 Sol (xhigh) has the lower benchmark cost at $1.156 vs $1.908. GPT-6.1 Sol (xhigh) is faster at 31.94s vs 84.15s, with pass rates of 97.0% vs 86.4%.

Last updated at: 2026-09-29

Compared models

Rank
#6
Total Output Tokens
99,719
Response Time (avg)
31.94s
Total Cost
$1.156
Rank
#24
Total Output Tokens
282,217
Response Time (avg)
84.15s
Total Cost
$1.908
Recommended model GPT-6.1 Sol (xhigh)

It has the best score here (9.9), while costing about 1.7x less than Grok 4.6 (high).

Detailed comparison

Metric GPT-6.1 Sol GPT-6.1 Sol xhigh Release: 2026-09-29 Grok 4.6 Grok 4.6 high Release: 2026-08-12
Score 9.9 9.3
Rank #6 #24
Reliability 10.0 10.0
Consistency 9.6 10.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 97.0% 86.4%
Flaky tests 1 0
Total Runs 66 66
Cost per result 5.504 10.042
Total Cost $1.156 $1.908
Input Price $2.000 / 1M $2.000 / 1M
Output Price $10.000 / 1M $6.000 / 1M
Total Input Tokens 79,232 107,275
Output Tokens 4,811 5,094
Reasoning Tokens 94,908 277,123
Response Time (avg) 31.94s 84.15s
Response Time (max) 491.85s 618.49s
Response Time (total) 702.66s 1851.26s
Parameters ~2T total (~150B active) ~1.7T total (~170B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#6 GPT-6.1 Sol

xhigh
Cost
$0.147
Time
279.4s
Tokens
14,790 tok

#24 SpaceXAI: Grok 4.6

high
Cost
$0.107
Time
265.1s
Tokens
18,001 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-6.1 Sol 10.0 10.0 100.0% 0 10.65s 7,302 393 4,483
Grok 4.6 10.0 10.0 100.0% 0 102.20s 9,579 379 51,829

Quick Compare

Switch Comparison Pair