Navigate
AI BENCHY
Advertise here

Compared models

GPT-5.6 Sol (low) vs Grok 4.6 (high) vs Claude Opus 5 (high) benchmark comparison: GPT-5.6 Sol (low) leads on Score with 9.5. GPT-5.6 Sol (low) leads on Reliability with 10.0. GPT-5.6 Sol (low) has the lowest Total Cost at $0.971. GPT-5.6 Sol (low) is fastest at 8.79s.

Last updated at: 2026-08-12

Rank
#5
Total Output Tokens
19,246
Response Time (avg)
8.79s
Total Cost
$0.971
Rank
#11
Total Output Tokens
265,604
Response Time (avg)
79.38s
Total Cost
$1.809
Rank
#10
Total Output Tokens
86,180
Response Time (avg)
20.44s
Total Cost
$2.835
Recommended model GPT-5.6 Sol (low)

It has the best score here (9.5), while costing about 2.4x less than the other models in this comparison.

Detailed comparison

Metric GPT-5.6 Sol GPT-5.6 Sol low Release: 2026-07-09 Grok 4.6 Grok 4.6 high Release: 2026-08-12 Claude Opus 5 Claude Opus 5 high Release: 2026-07-25
Score 9.5 9.2 9.2
Rank #5 #11 #10
Reliability 10.0 10.0 10.0
Consistency 9.2 9.6 9.6
Attempts 66/66 66/66 66/66
Tests Correct
Attempt pass rate 86.4% 84.9% 84.9%
Flaky tests 2 1 1
Total Runs 66 66 66
Cost per result 5.391 10.046 15.747
Total Cost $0.971 $1.809 $2.835
Input Price $5.000 / 1M $2.000 / 1M $5.000 / 1M
Output Price $30.000 / 1M $6.000 / 1M $25.000 / 1M
Total Input Tokens 78,571 107,266 135,959
Output Tokens 4,476 5,094 67,066
Reasoning Tokens 14,770 260,510 19,114
Response Time (avg) 8.79s 79.38s 20.44s
Response Time (max) 53.91s 618.49s 159.01s
Response Time (total) 193.33s 1746.29s 449.68s
Parameters ~2T total (~150B active) ~1.7T total (~170B active) ~5T total (~500B active)
Availability Closed Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#5 GPT-5.6 Sol

low
Cost
$0.062
Time
26.7s
Tokens
2,150 tok

#11 SpaceXAI: Grok 4.6

high
Cost
$0.107
Time
265.1s
Tokens
18,001 tok

#10 Claude Opus 5

high
Cost
$0.265
Time
144.5s
Tokens
10,741 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Sol 10.0 10.0 100.0% 0 11.25s 7,302 412 3,482
Grok 4.6 10.0 10.0 100.0% 0 102.20s 9,579 379 51,829
Claude Opus 5 10.0 10.0 100.0% 0 12.73s 10,590 7,547 1,721

Quick Compare

Switch Comparison Pair