Navigate
Advertise here

GPT-5.6 Luna (medium) vs Grok 4.7 (low)

The average score is effectively tied at 7.4 vs 7.4. GPT-5.6 Luna (medium) has the lower benchmark cost at $0.101 vs $2.263. GPT-5.6 Luna (medium) is faster at 9.64s vs 63.19s, with pass rates of 62.3% vs 73.9%.

Last updated at: 2026-10-01

Compared models

Rank
#122
Total Output Tokens
47,986
Response Time (avg)
9.64s
Total Cost
$0.101
Rank
#127
Total Output Tokens
255,329
Response Time (avg)
63.19s
Total Cost
$2.263
Recommended model GPT-5.6 Luna (medium)

It has the best score here (7.4), while costing about 22.6x less than Grok 4.7 (low).

Detailed comparison

Metric GPT-5.6 Luna GPT-5.6 Luna medium Release: 2026-07-09 Grok 4.7 Grok 4.7 low Release: 2026-09-21
Score 7.4 7.4
Rank #122 #127
Reliability 10.0 10.0
Consistency 9.0 8.3
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 62.3% 73.9%
Flaky tests 3 5
Total Runs 69 69
Cost per result 2.822 13.804
Total Cost $0.101 $2.263
Input Price $0.200 / 1M $2.000 / 1M
Output Price $1.200 / 1M $6.000 / 1M
Total Input Tokens 212,780 365,119
Output Tokens 6,849 8,118
Reasoning Tokens 41,137 247,211
Response Time (avg) 9.64s 63.19s
Response Time (max) 58.09s 389.63s
Response Time (total) 221.62s 1453.28s
Parameters ~400B total (~17B active) ~1.7T total (~170B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#122 GPT-5.6 Luna

medium
Cost
$0.014
Time
14.6s
Tokens
2,362 tok

#127 SpaceXAI: Grok 4.7

low
Cost
$0.007
Time
15.2s
Tokens
1,537 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Luna 5.4 7.2 44.4% 1 10.38s 7,302 564 9,928
Grok 4.7 8.4 7.4 88.9% 1 189.08s 9,579 346 99,053

Quick Compare

Switch Comparison Pair