Navigate
Advertise here

Ling 3.0 Tiny (medium) vs GPT-6 Luna Decisions (default)

Ling 3.0 Tiny (medium) leads on average score with 3.6 vs 2.4. Ling 3.0 Tiny (medium) has the lower benchmark cost at $0.000 vs $0.005. GPT-6 Luna Decisions (default) is faster at 495ms vs 64.99s, with pass rates of 23.2% vs 13.0%.

Last updated at: 2026-10-07

Compared models

Rank
#370
Total Output Tokens
911,748
Response Time (avg)
64.99s
Total Cost
$0.000
Rank
#390
Total Output Tokens
0
Response Time (avg)
495ms
Total Cost
$0.005
Recommended model Ling 3.0 Tiny (medium)

It has the strongest score in this comparison (3.6) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric Ling 3.0 Tiny Ling 3.0 Tiny medium Release: 2026-08-07 GPT-6 Luna Decisions GPT-6 Luna Decisions default Release: 2026-10-07
Score 3.6 2.4
Rank #370 #390
Reliability 9.6 10.0
Consistency 8.9 5.2
Attempts 69/69 36/69
Tests Correct
Attempt pass rate 23.2% 13.0%
Flaky tests 3 0
Total Runs 69 36
Cost per result 0.000 0.159
Total Cost $0.000 $0.005
Input Price $0.000 / 1M $0.100 / 1M
Output Price $0.000 / 1M $0.000 / 1M
Cache Read Price N/A $0.000 / 1M
Cache Write Price N/A $0.000 / 1M
Total Input Tokens 103,898 47,583
Output Tokens 148,304 0
Reasoning Tokens 774,972 0
Response Time (avg) 64.99s 495ms
Response Time (max) 262.20s 842ms
Response Time (total) 1429.68s 5.94s
Parameters 7.9B total (1.3B active) ~400B total (~17B active)
Availability Open source Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#370 Ling 3.0 Tiny

medium
No output was saved. The original provider response or failure reason is unavailable.
Cost
$0.000
Time
177.4s
Tokens
6,873 tok

#390 GPT-6 Luna Decisions

default
No showcase result has been generated for this model yet.
Cost
N/A
Time
-
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Ling 3.0 Tiny 2.9 10.0 0.0% 0 211.99s 7,948 70,838 206,006
GPT-6 Luna Decisions 2.5 6.7 0.0% 0 447ms 22,197 0 0

Quick Compare

Switch Comparison Pair