Navigate
AI BENCHY
Advertise here

GPT-5.4 (medium) vs Step 3.7 Flash (low)

GPT-5.4 (medium) leads on average score with 8.5 vs 7.3. Step 3.7 Flash (low) has the lower benchmark cost at $0.454 vs $1.533. Step 3.7 Flash (low) is faster at 20.68s vs 23.10s, with pass rates of 77.3% vs 68.2%.

Last updated at: 2026-07-28

Rank
#24
Total Output Tokens
88,670
Response Time (avg)
23.10s
Total Cost
$1.533
Rank
#77
Total Output Tokens
376,581
Response Time (avg)
20.68s
Total Cost
$0.454
Recommended model GPT-5.4 (medium)

It has the strongest score in this comparison (8.5) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric GPT-5.4 GPT-5.4 medium Release: 2026-03-05 Step 3.7 Flash Step 3.7 Flash low Release: 2026-05-29
Score 8.5 7.3
Rank #24 #77
Reliability 10.0 10.0
Consistency 8.6 8.1
Tests Correct
Attempt pass rate 77.3% 68.2%
Flaky tests 4 5
Total Runs 66 66
Cost per result 10.220 3.782
Total Cost $1.533 $0.454
Input Price $2.500 / 1M $0.200 / 1M
Output Price $15.000 / 1M $1.150 / 1M
Total Input Tokens 81,127 103,833
Output Tokens 6,155 376,581
Reasoning Tokens 82,515 0
Response Time (avg) 23.10s 20.68s
Response Time (max) 100.41s 124.75s
Response Time (total) 508.26s 455.01s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#24 GPT-5.4

medium
Cost
$0.214
Time
199.6s
Tokens
14,349 tok

#77 Step 3.7 Flash

low
Invalid SVG
Cost
$0.004
Time
25.3s
Tokens
3,072 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 8.8 7.8 88.9% 1 44.36s 7,305 433 24,216
Step 3.7 Flash 8.2 7.2 88.9% 1 9.46s 7,437 18,685 0

Quick Compare

Switch Comparison Pair