Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

OpenAI: GPT-5.4 Mini vs StepFun: Step 3.7 Flash

Summary

GPT-5.4 Mini vs Step 3.7 Flash benchmark comparison: Step 3.7 Flash leads on average score with 8.5 vs 8.0. Step 3.7 Flash has the lower benchmark cost at $0.376 vs $0.526. Step 3.7 Flash is faster at 20.35s vs 22.34s, with pass rates of 73.0% vs 73.0%.

Recommended model: Step 3.7 Flash - It has the strongest score in this comparison (8.5) and the best overall balance of cost and response time across all 2 models.

Last updated at: 2026-06-12

Metric GPT-5.4 Mini GPT-5.4 Mini medium Release: 2026-03-17 Step 3.7 Flash Step 3.7 Flash medium Release: 2026-05-29
Score 8.0 8.5
Rank #30 #23
Reliability 10.0 9.9
Consistency 8.0 9.3
Tests Correct
Attempt pass rate 73.0% 73.0%
Flaky tests 5 2
Total Runs 63 61
Cost per result 4.381 2.686
Total Cost $0.526 $0.376
Input Price $0.750 / 1M $0.200 / 1M
Output Price $4.500 / 1M $1.150 / 1M
Total Input Tokens 34,116 39,981
Output Tokens 2,181 319,958
Reasoning Tokens 108,937 0
Response Time (avg) 22.34s 20.35s
Response Time (max) 138.75s 113.98s
Response Time (total) 469.20s 427.42s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#30 GPT-5.4 Mini

medium
Cost
$0.056
Time
95.5s
Tokens
12,464 tok

#23 Step 3.7 Flash

medium
Cost
$0.006
Time
46.2s
Tokens
4,466 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 8.6 7.9 91.7% 1 4.05s 606 296 2,876
Step 3.7 Flash 8.7 7.9 91.7% 1 9.65s 756 32,185 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 8.4 7.4 88.9% 1 57.87s 7,305 467 40,902
Step 3.7 Flash 8.8 7.8 88.9% 1 27.42s 7,437 44,797 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 10.0 10.0 100.0% 0 17.81s 11,019 317 4,317
Step 3.7 Flash 10.0 10.0 100.0% 0 9.06s 13,683 7,106 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 10.0 10.0 100.0% 0 2.43s 7,140 234 650
Step 3.7 Flash 10.0 10.0 100.0% 0 2.75s 7,398 3,020 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 4.1 4.4 44.5% 2 65.31s 619 60 43,286
Step 3.7 Flash 7.7 10.0 66.7% 0 48.27s 708 70,347 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 4.5 10.0 0.0% 0 3.72s 477 150 510
Step 3.7 Flash 4.0 10.0 0.0% 0 6.85s 525 3,987 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 9.8 10.0 100.0% 0 2.13s 660 96 1,185
Step 3.7 Flash 9.8 10.0 100.0% 0 1.83s 735 2,166 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 7.8 10.0 66.7% 0 4.37s 642 278 2,443
Step 3.7 Flash 5.7 9.9 33.3% 0 6.19s 756 15,071 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 4.7 1.6 66.7% 1 9.62s 5,453 251 2,594
Step 3.7 Flash 10.0 10.0 100.0% 0 4.16s 7,746 2,115 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 3.0 10.0 0.0% 0 30.10s 195 32 10,174
Step 3.7 Flash 3.0 10.0 0.0% 0 113.98s 237 139,164 0

Quick Compare

Switch Comparison Pair