Navigate
Advertise here

GPT-6 Luna Decisions (default) vs Tev1 4B Experimental (default)

GPT-6 Luna Decisions (default) leads on average score with 2.4 vs 1.0. Tev1 4B Experimental (default) has the lower benchmark cost at $0.002 vs $0.005. GPT-6 Luna Decisions (default) is faster at 495ms vs 498ms, with pass rates of 13.0% vs 13.0%.

Last updated at: 2026-10-07

Compared models

Rank
#390
Total Output Tokens
0
Response Time (avg)
495ms
Total Cost
$0.005
Rank
#397
Total Output Tokens
96
Response Time (avg)
498ms
Total Cost
$0.002
Recommended model GPT-6 Luna Decisions (default)

It has the strongest score in this comparison (2.4) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric GPT-6 Luna Decisions GPT-6 Luna Decisions default Release: 2026-10-07 Tev1 4B Experimental Tev1 4B Experimental default Release: 2026-09-30
Score 2.4 1.0
Rank #390 #397
Reliability 10.0 10.0
Consistency 5.2 3.0
Attempts 36/69 21/69
Tests Correct
Attempt pass rate 13.0% 13.0%
Flaky tests 0 0
Total Runs 36 21
Cost per result 0.159 0.048
Total Cost $0.005 $0.002
Input Price $0.100 / 1M $0.042 / 1M
Output Price $0.000 / 1M $0.000 / 1M
Cache Read Price $0.000 / 1M N/A
Cache Write Price $0.000 / 1M N/A
Total Input Tokens 47,583 33,864
Output Tokens 0 96
Reasoning Tokens 0 0
Response Time (avg) 495ms 498ms
Response Time (max) 842ms 602ms
Response Time (total) 5.94s 3.49s
Parameters ~400B total (~17B active) 4B
Availability Closed Weights available

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-6 Luna Decisions 2.5 6.7 0.0% 0 447ms 22,197 0 0
Tev1 4B Experimental 0.0 0.0 0.0% 0 0ms 0 0 0

Quick Compare

Switch Comparison Pair

Command A+lowvsGPT-6 Luna DecisionsdefaultCommand A+nonevsGPT-6 Luna DecisionsdefaultGemini 3 Flash PreviewhighvsGPT-6 Luna DecisionsdefaultCommand A+highvsGPT-6 Luna DecisionsdefaultGemini 3 Flash PreviewhighvsTev1 4B ExperimentaldefaultCommand A+mediumvsGPT-6 Luna DecisionsdefaultGPT-6 Luna DecisionsdefaultvsGrok 4.20noneLing 3.0 TinymediumvsGPT-6 Luna DecisionsdefaultLing 3.0 TinylowvsGPT-6 Luna DecisionsdefaultLing 3.0 TinyhighvsGPT-6 Luna DecisionsdefaultNemotron 3.5 LightningnoneFree AvailablevsGPT-6 Luna DecisionsdefaultGPT-6 Luna DecisionsdefaultvsQwen3.5-9Bmedium