Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Anthropic: Claude Opus 4.8 vs StepFun: Step 3.7 Flash

Summary

Claude Opus 4.8 vs Step 3.7 Flash benchmark comparison: Claude Opus 4.8 leads on average score with 8.8 vs 7.7. Step 3.7 Flash has the lower benchmark cost at $0.341 vs $1.107. Claude Opus 4.8 is faster at 9.66s vs 15.74s, with pass rates of 84.1% vs 68.3%.

Recommended model: Step 3.7 Flash - It offers the best overall trade-off: a competitive score (7.7), lower cost than Claude Opus 4.8, and balanced response time.

Last updated at: 2026-06-17

Metric Claude Opus 4.8 Claude Opus 4.8 medium Release: 2026-05-28 Step 3.7 Flash Step 3.7 Flash low Release: 2026-05-29
Score 8.8 7.7
Rank #12 #39
Reliability 10.0 10.0
Consistency 9.6 8.4
Tests Correct
Attempt pass rate 84.1% 68.3%
Flaky tests 1 4
Total Runs 63 63
Cost per result 6.512 2.840
Total Cost $1.107 $0.341
Input Price $5.000 / 1M $0.200 / 1M
Output Price $25.000 / 1M $1.150 / 1M
Total Input Tokens 61,007 40,101
Output Tokens 26,495 289,325
Reasoning Tokens 5,901 0
Response Time (avg) 9.66s 15.74s
Response Time (max) 38.03s 124.75s
Response Time (total) 202.89s 330.63s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#12 Claude Opus 4.8

medium
Cost
$0.057
Time
23.1s
Tokens
2,412 tok

#39 Step 3.7 Flash

low
Invalid SVG
Cost
$0.004
Time
25.3s
Tokens
3,072 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 834 1,179 478
Step 3.7 Flash 8.7 7.9 91.7% 1 4.02s 756 10,896 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 15.33s 10,590 9,945 1,381
Step 3.7 Flash 8.2 7.2 88.9% 1 9.46s 7,437 18,685 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 9.8 10.0 100.0% 0 38.03s 23,561 5,260 1,588
Step 3.7 Flash 10.0 10.0 100.0% 0 7.98s 13,683 6,426 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 7.1 5.6 83.3% 1 12.29s 10,503 481 312
Step 3.7 Flash 7.3 5.8 83.3% 1 2.29s 7,398 2,667 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 5.3 10.0 33.3% 0 14.15s 975 7,477 900
Step 3.7 Flash 5.3 7.2 44.4% 1 43.31s 828 104,487 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 2.46s 708 237 0
Step 3.7 Flash 3.4 9.3 0.0% 0 7.00s 525 4,604 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.32s 909 373 320
Step 3.7 Flash 9.8 10.0 100.0% 0 1.58s 735 1,857 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 894 791 483
Step 3.7 Flash 5.5 9.9 33.3% 0 1.84s 756 3,564 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 8.96s 11,775 301 225
Step 3.7 Flash 10.0 10.0 100.0% 0 3.25s 7,746 1,360 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 3.0 10.0 0.0% 0 6.14s 258 451 214
Step 3.7 Flash 3.0 10.0 0.0% 0 124.75s 237 134,779 0

Quick Compare

Switch Comparison Pair