Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

GPT-5.6 Luna vs Inkling Small

The average score is effectively tied at 5.4 vs 5.3. GPT-5.6 Luna has the lower benchmark cost at $0.015 vs $0.058. GPT-5.6 Luna is faster at 1.50s vs 1.51s, with pass rates of 34.9% vs 24.2%.

Last updated at: 2026-08-01

Rank
#184
Total Output Tokens
6,709
Response Time (avg)
1.50s
Total Cost
$0.015
Rank
#187
Total Output Tokens
8,770
Response Time (avg)
1.51s
Total Cost
$0.058
Recommended model GPT-5.6 Luna

It has the best score here (5.4), while costing about 4.1x less than Inkling Small.

Detailed comparison

Metric GPT-5.6 Luna GPT-5.6 Luna none Release: 2026-07-09 Inkling Small Inkling Small none Release: 2026-08-01
Score 5.4 5.3
Rank #184 #187
Reliability 10.0 10.0
Consistency 8.8 9.6
Tests Correct
Attempt pass rate 34.9% 24.2%
Flaky tests 3 1
Total Runs 66 66
Cost per result 2.360 1.158
Total Cost $0.015 $0.058
Input Price $0.100 / 1M $0.500 / 1M
Output Price $0.601 / 1M $1.200 / 1M
Total Input Tokens 101,323 94,724
Output Tokens 6,709 8,770
Reasoning Tokens 0 0
Response Time (avg) 1.50s 1.51s
Response Time (max) 10.57s 12.35s
Response Time (total) 32.91s 33.27s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#184 GPT-5.6 Luna

none
Cost
$0.016
Time
15.8s
Tokens
2,685 tok

#187 Thinking Machines: Inkling Small

none
Cost
$0.008
Time
42.1s
Tokens
6,319 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Luna 3.8 7.2 22.2% 1 980ms 7,302 459 0
Inkling Small 3.6 10.0 0.0% 0 759ms 7,356 454 0

Quick Compare

Switch Comparison Pair