Navigate
AI BENCHY
Advertise here

GPT-5.6 Sol vs Grok 4.20 (medium)

Grok 4.20 (medium) leads on average score with 7.1 vs 6.9. GPT-5.6 Sol has the lower benchmark cost at $0.524 vs $0.777. GPT-5.6 Sol is faster at 2.16s vs 29.47s, with pass rates of 59.1% vs 63.6%.

Last updated at: 2026-07-25

Rank
#91
Total Output Tokens
4,357
Response Time (avg)
2.16s
Total Cost
$0.524
Rank
#83
Total Output Tokens
259,340
Response Time (avg)
29.47s
Total Cost
$0.777
Recommended model GPT-5.6 Sol

Its score stays close to the best score here (6.9 vs 7.1), while responding about 13.6x faster than Grok 4.20 (medium).

Detailed comparison

Metric GPT-5.6 Sol GPT-5.6 Sol none Release: 2026-07-09 Grok 4.20 Grok 4.20 medium Release: 2026-03-31
Score 6.9 7.1
Rank #91 #83
Reliability 10.0 10.0
Consistency 9.0 8.5
Tests Correct
Attempt pass rate 59.1% 63.6%
Flaky tests 3 4
Total Runs 66 66
Cost per result 4.761 9.709
Total Cost $0.524 $0.777
Input Price $5.000 / 1M $1.250 / 1M
Output Price $30.000 / 1M $2.500 / 1M
Total Input Tokens 78,593 102,791
Output Tokens 4,357 5,363
Reasoning Tokens 0 253,977
Response Time (avg) 2.16s 29.47s
Response Time (max) 12.81s 199.66s
Response Time (total) 47.62s 648.35s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#91 GPT-5.6 Sol

none
Cost
$0.116
Time
78.7s
Tokens
3,944 tok

#83 xAI: Grok 4.20

medium
Cost
$0.041
Time
110.3s
Tokens
16,336 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Sol 5.5 10.0 33.3% 0 1.39s 7,302 390 0
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150

Quick Compare

Switch Comparison Pair