Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Qwen: Qwen3.5 Plus 2026-02-15 vs StepFun: Step 3.7 Flash

Summary

Qwen3.5 Plus 2026-02-15 vs Step 3.7 Flash benchmark comparison: Step 3.7 Flash leads on average score with 7.1 vs 5.8. Qwen3.5 Plus 2026-02-15 has the lower benchmark cost at $0.016 vs $1.148. Qwen3.5 Plus 2026-02-15 is faster at 2.31s vs 64.46s, with pass rates of 46.0% vs 63.5%.

Recommended model: Qwen3.5 Plus 2026-02-15 - It offers the best overall trade-off: a competitive score (5.8), lower cost than Step 3.7 Flash, and balanced response time.

Last updated at: 2026-06-18

Metric Qwen3.5 Plus 2026-02-15 Qwen3.5 Plus 2026-02-15 none Release: 2026-02-15 Step 3.7 Flash Step 3.7 Flash high Release: 2026-05-29
Score 5.8 7.1
Rank #106 #63
Reliability 10.0 10.0
Consistency 9.4 8.2
Tests Correct
Attempt pass rate 46.0% 63.5%
Flaky tests 2 4
Total Runs 63 63
Cost per result 0.204 10.434
Total Cost $0.016 $1.148
Input Price $0.260 / 1M $0.200 / 1M
Output Price $1.560 / 1M $1.150 / 1M
Total Input Tokens 45,864 38,391
Output Tokens 2,480 991,355
Reasoning Tokens 0 0
Response Time (avg) 2.31s 64.46s
Response Time (max) 6.65s 364.99s
Response Time (total) 34.63s 1353.57s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#106 Qwen3.5 Plus 2026-02-15

none
Cost
$0.012
Time
153.2s
Tokens
7,787 tok

#63 Step 3.7 Flash

high
Cost
$0.007
Time
63.6s
Tokens
6,030 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 4.8 10.0 25.0% 0 1.91s 696 517 0
Step 3.7 Flash 10.0 10.0 100.0% 0 13.40s 696 42,656 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 4.3 7.9 11.1% 1 2.05s 7,913 473 0
Step 3.7 Flash 4.0 6.0 22.2% 1 206.21s 6,057 327,340 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 3.0 10.0 0.0% 0 6.65s 18,304 314 0
Step 3.7 Flash 10.0 10.0 100.0% 0 13.01s 13,638 8,802 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 1.89s 7,794 243 0
Step 3.7 Flash 10.0 10.0 100.0% 0 14.72s 7,368 23,113 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 5.3 10.0 33.3% 0 1.17s 789 17 0
Step 3.7 Flash 4.1 4.4 44.5% 2 149.64s 783 410,502 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 4.4 3.0 33.3% 1 2.26s 522 117 0
Step 3.7 Flash 5.5 10.0 0.0% 0 4.17s 510 2,862 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 1.67s 711 72 0
Step 3.7 Flash 9.8 10.0 100.0% 0 1.52s 705 2,010 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 7.7 10.0 66.7% 0 2.71s 714 494 0
Step 3.7 Flash 5.3 7.2 44.4% 1 10.22s 711 25,422 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 3.33s 8,211 222 0
Step 3.7 Flash 10.0 10.0 100.0% 0 2.79s 7,701 1,172 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5 Plus 2026-02-15 3.0 10.0 0.0% 0 1.11s 210 11 0
Step 3.7 Flash 3.0 10.0 0.0% 0 149.34s 222 147,476 0

Quick Compare

Switch Comparison Pair