Navigate
Advertise here

GPT-6.1 Sol (max) vs Grok 4.6 (high)

GPT-6.1 Sol (max) leads on average score with 9.9 vs 9.3. GPT-6.1 Sol (max) has the lower benchmark cost at $0.740 vs $1.908. GPT-6.1 Sol (max) is faster at 61.16s vs 84.15s, with pass rates of 95.5% vs 86.4%.

Last updated at: 2026-09-29

Compared models

Rank
#4
Total Output Tokens
58,145
Response Time (avg)
61.16s
Total Cost
$0.740
Rank
#24
Total Output Tokens
282,217
Response Time (avg)
84.15s
Total Cost
$1.908
Recommended model GPT-6.1 Sol (max)

It has the best score here (9.9), while costing about 2.6x less than Grok 4.6 (high).

Detailed comparison

Metric GPT-6.1 Sol GPT-6.1 Sol max Release: 2026-09-29 Grok 4.6 Grok 4.6 high Release: 2026-08-12
Score 9.9 9.3
Rank #4 #24
Reliability 9.6 10.0
Consistency 10.0 10.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 95.5% 86.4%
Flaky tests 0 0
Total Runs 66 66
Cost per result 3.523 10.042
Total Cost $0.740 $1.908
Input Price $2.000 / 1M $2.000 / 1M
Output Price $10.000 / 1M $6.000 / 1M
Total Input Tokens 79,136 107,275
Output Tokens 4,913 5,094
Reasoning Tokens 53,232 277,123
Response Time (avg) 61.16s 84.15s
Response Time (max) 960.01s 618.49s
Response Time (total) 1345.41s 1851.26s
Parameters ~2T total (~150B active) ~1.7T total (~170B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#4 GPT-6.1 Sol

max
Cost
$0.272
Time
584.0s
Tokens
27,273 tok

#24 SpaceXAI: Grok 4.6

high
Cost
$0.107
Time
265.1s
Tokens
18,001 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-6.1 Sol 10.0 10.0 100.0% 0 13.57s 7,302 398 7,668
Grok 4.6 10.0 10.0 100.0% 0 102.20s 9,579 379 51,829

Quick Compare

Switch Comparison Pair