Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

GPT-5.6 Sol (low) vs Grok 4.6 (high) vs Claude Opus 5 (high) benchmark comparison: GPT-5.6 Sol (low) leads on Score with 9.5. GPT-5.6 Sol (low) leads on Reliability with 10.0. GPT-5.6 Sol (low) has the lowest Total Cost at $0.357. GPT-5.6 Sol (low) is fastest at 9.04s.

Last updated at: 2026-09-24

Compared models

Rank
#9
Total Output Tokens
19,946
Response Time (avg)
9.04s
Total Cost
$0.357
Rank
#19
Total Output Tokens
282,217
Response Time (avg)
84.15s
Total Cost
$1.908
Rank
#24
Total Output Tokens
95,583
Response Time (avg)
22.51s
Total Cost
$3.070
Recommended model GPT-5.6 Sol (low)

It has the best score here (9.5), while costing about 7.0x less than the other models in this comparison.

Detailed comparison

Metric GPT-5.6 Sol GPT-5.6 Sol low Release: 2026-07-09 Grok 4.6 Grok 4.6 high Release: 2026-08-12 Claude Opus 5 Claude Opus 5 high Release: 2026-07-25
Score 9.5 9.3 9.2
Rank #9 #19 #24
Reliability 10.0 10.0 10.0
Consistency 9.2 10.0 10.0
Attempts 66/66 66/66 66/66
Tests Correct
Attempt pass rate 86.4% 86.4% 81.8%
Flaky tests 2 0 0
Total Runs 66 66 66
Cost per result 5.508 10.042 17.052
Total Cost $0.357 $1.908 $3.070
Input Price $2.000 / 1M $2.000 / 1M $5.000 / 1M
Output Price $10.000 / 1M $6.000 / 1M $25.000 / 1M
Total Input Tokens 78,580 107,275 135,956
Output Tokens 4,476 5,094 64,624
Reasoning Tokens 15,470 277,123 30,959
Response Time (avg) 9.04s 84.15s 22.51s
Response Time (max) 53.91s 618.49s 159.01s
Response Time (total) 198.97s 1851.26s 495.12s
Parameters ~2T total (~150B active) ~1.7T total (~170B active) ~5T total (~500B active)
Availability Closed Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#9 GPT-5.6 Sol

low
Cost
$0.062
Time
26.7s
Tokens
2,150 tok

#19 SpaceXAI: Grok 4.6

high
Cost
$0.107
Time
265.1s
Tokens
18,001 tok

#24 Claude Opus 5

high
Cost
$0.265
Time
144.5s
Tokens
10,741 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Sol 10.0 10.0 100.0% 0 11.25s 7,302 412 3,482
Grok 4.6 10.0 10.0 100.0% 0 102.20s 9,579 379 51,829
Claude Opus 5 10.0 10.0 100.0% 0 12.73s 10,590 7,547 1,721

Quick Compare

Switch Comparison Pair