Navigate
AI BENCHY
Advertise here

Qwen3.7 Flash (medium) vs Step 3.7 Flash (high)

Qwen3.7 Flash (medium) leads on average score with 7.0 vs 6.9. Qwen3.7 Flash (medium) has the lower benchmark cost at $0.064 vs $1.207. Qwen3.7 Flash (medium) is faster at 45.90s vs 64.68s, with pass rates of 65.2% vs 63.6%.

Last updated at: 2026-07-28

Rank
#91
Total Output Tokens
462,623
Response Time (avg)
45.90s
Total Cost
$0.064
Rank
#97
Total Output Tokens
1,032,395
Response Time (avg)
64.68s
Total Cost
$1.207
Recommended model Qwen3.7 Flash (medium)

It has the best score here (7.0), while costing about 19.0x less than Step 3.7 Flash (high).

Detailed comparison

Metric Qwen3.7 Flash Qwen3.7 Flash medium Release: 2026-07-28 Step 3.7 Flash Step 3.7 Flash high Release: 2026-05-29
Score 7.0 6.9
Rank #91 #97
Reliability 10.0 10.0
Consistency 7.5 8.0
Tests Correct
Attempt pass rate 65.2% 63.6%
Flaky tests 7 5
Total Runs 66 66
Cost per result 0.636 10.973
Total Cost $0.064 $1.207
Input Price $0.030 / 1M $0.200 / 1M
Output Price $0.130 / 1M $1.150 / 1M
Total Input Tokens 114,468 98,691
Output Tokens 12,786 1,032,395
Reasoning Tokens 449,837 0
Response Time (avg) 45.90s 64.68s
Response Time (max) 593.64s 364.99s
Response Time (total) 1009.84s 1423.01s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#91 Qwen3.7 Flash

medium
Invalid SVG
Cost
$0.000
Time
600.0s
Tokens
0 tok

#97 Step 3.7 Flash

high
Cost
$0.007
Time
63.6s
Tokens
6,030 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.7 Flash 6.2 6.9 55.6% 1 46.78s 7,893 563 61,902
Step 3.7 Flash 4.0 6.0 22.2% 1 206.21s 6,057 327,340 0

Quick Compare

Switch Comparison Pair