Navigate
AI BENCHY
Advertise here

Grok 4.5 (low) vs Grok 4.6 (medium)

Grok 4.5 (low) leads on average score with 8.4 vs 8.3. Grok 4.5 (low) has the lower benchmark cost at $0.935 vs $1.656. Grok 4.5 (low) is faster at 15.56s vs 71.86s, with pass rates of 75.8% vs 80.3%.

Last updated at: 2026-08-12

Rank
#35
Total Output Tokens
113,951
Response Time (avg)
15.56s
Total Cost
$0.935
Rank
#40
Total Output Tokens
240,338
Response Time (avg)
71.86s
Total Cost
$1.656
Recommended model Grok 4.5 (low)

It has the best score here (8.4), while costing about 1.8x less than Grok 4.6 (medium).

Detailed comparison

Metric Grok 4.5 Grok 4.5 low Release: 2026-07-08 Grok 4.6 Grok 4.6 medium Release: 2026-08-12
Score 8.4 8.3
Rank #35 #40
Reliability 10.0 10.0
Consistency 9.7 8.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 75.8% 80.3%
Flaky tests 1 5
Total Runs 66 66
Cost per result 5.844 11.038
Total Cost $0.935 $1.656
Input Price $2.000 / 1M $2.000 / 1M
Output Price $6.000 / 1M $6.000 / 1M
Total Input Tokens 125,596 106,778
Output Tokens 7,505 4,811
Reasoning Tokens 106,446 235,527
Response Time (avg) 15.56s 71.86s
Response Time (max) 205.28s 514.35s
Response Time (total) 342.32s 1580.96s
Parameters ~1.7T total (~170B active) ~1.7T total (~170B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#35 xAI: Grok 4.5

low
Cost
$0.011
Time
12.4s
Tokens
2,018 tok

#40 SpaceXAI: Grok 4.6

medium
Cost
$0.028
Time
72.1s
Tokens
4,771 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Grok 4.5 10.0 10.0 100.0% 0 13.72s 9,579 303 15,641
Grok 4.6 8.2 7.2 88.9% 1 120.05s 9,579 351 64,226

Quick Compare

Switch Comparison Pair