Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.6 vs StepFun: Step 3.7 Flash

Summary

Claude Opus 4.6 vs Step 3.7 Flash benchmark comparison: Step 3.7 Flash leads on average score with 7.3 vs 7.0. Step 3.7 Flash has the lower benchmark cost at $0.341 vs $2.053. Step 3.7 Flash is faster at 15.74s vs 25.89s, with pass rates of 61.9% vs 68.3%.

Recommended model: Step 3.7 Flash - It has the best score here (7.3), while costing about 6.0x less than Claude Opus 4.6.

Last updated at: 2026-06-04

Metric Claude Opus 4.6 Claude Opus 4.6 medium Release: 2026-02-05 Step 3.7 Flash Step 3.7 Flash low Release: 2026-05-29
Score 7.0 7.3
Rank #69 #57
Reliability 10.0 10.0
Consistency 8.8 8.4
Tests Correct
Attempt pass rate 61.9% 68.3%
Flaky tests 3 4
Total Runs 63 63
Cost per result 17.103 2.840
Total Cost $2.053 $0.341
Input Price $5.000 / 1M $0.200 / 1M
Output Price $25.000 / 1M $1.150 / 1M
Total Input Tokens 53,227 40,101
Output Tokens 47,446 289,325
Reasoning Tokens 24,000 0
Response Time (avg) 25.89s 15.74s
Response Time (max) 83.40s 124.75s
Response Time (total) 362.49s 330.63s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#69 Claude Opus 4.6

medium
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

#57 Step 3.7 Flash

low
Invalid SVG
Cost
$0.004
Time
25.3s
Tokens
3,072 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 6.4 5.8 66.7% 2 7.45s 840 986 1,071
Step 3.7 Flash 8.7 7.9 91.7% 1 4.02s 756 10,896 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 5.7 7.1 44.4% 1 30.10s 8,522 13,057 4,121
Step 3.7 Flash 8.2 7.2 88.9% 1 9.46s 7,437 18,685 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 76.66s 20,685 8,178 5,194
Step 3.7 Flash 10.0 10.0 100.0% 0 7.98s 13,683 6,426 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 7.37s 8,676 691 757
Step 3.7 Flash 7.3 5.8 83.3% 1 2.29s 7,398 2,667 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 3.0 10.0 0.0% 0 83.40s 674 14,642 8,687
Step 3.7 Flash 5.3 7.2 44.4% 1 43.31s 828 104,487 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 5.04s 564 188 292
Step 3.7 Flash 3.4 9.3 0.0% 0 7.00s 525 4,604 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 2.43s 792 266 467
Step 3.7 Flash 9.8 10.0 100.0% 0 1.58s 735 1,857 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 7.7 10.0 66.7% 0 4.71s 816 532 630
Step 3.7 Flash 5.5 9.9 33.3% 0 1.84s 756 3,564 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 9.73s 11,454 861 329
Step 3.7 Flash 10.0 10.0 100.0% 0 3.25s 7,746 1,360 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.6 3.0 10.0 0.0% 0 63.24s 204 8,045 2,452
Step 3.7 Flash 3.0 10.0 0.0% 0 124.75s 237 134,779 0

Quick Compare

Switch Comparison Pair