Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.7 vs StepFun: Step 3.7 Flash

Summary

Claude Opus 4.7 vs Step 3.7 Flash benchmark comparison: Claude Opus 4.7 leads on average score with 8.7 vs 7.7. Step 3.7 Flash has the lower benchmark cost at $0.341 vs $0.679. Claude Opus 4.7 is faster at 4.73s vs 15.74s, with pass rates of 82.5% vs 68.3%.

Recommended model: Claude Opus 4.7 - It has the best score here (8.7), while responding about 3.3x faster than Step 3.7 Flash.

Last updated at: 2026-06-12

Metric Claude Opus 4.7 Claude Opus 4.7 medium Release: 2026-04-16 Step 3.7 Flash Step 3.7 Flash low Release: 2026-05-29
Score 8.7 7.7
Rank #17 #42
Reliability 10.0 10.0
Consistency 9.6 8.4
Tests Correct
Attempt pass rate 82.5% 68.3%
Flaky tests 1 4
Total Runs 63 63
Cost per result 3.991 2.840
Total Cost $0.679 $0.341
Input Price $5.000 / 1M $0.200 / 1M
Output Price $25.000 / 1M $1.150 / 1M
Total Input Tokens 65,406 40,101
Output Tokens 11,858 289,325
Reasoning Tokens 2,198 0
Response Time (avg) 4.73s 15.74s
Response Time (max) 23.18s 124.75s
Response Time (total) 94.51s 330.63s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#17 Claude Opus 4.7

medium
Cost
$0.059
Time
26.8s
Tokens
2,475 tok

#42 Step 3.7 Flash

low
Invalid SVG
Cost
$0.004
Time
25.3s
Tokens
3,072 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 8.3 10.0 75.0% 0 1.85s 894 348 0
Step 3.7 Flash 8.7 7.9 91.7% 1 4.02s 756 10,896 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 7.6 7.2 77.8% 1 12.96s 10,635 7,629 1,114
Step 3.7 Flash 8.2 7.2 88.9% 1 9.46s 7,437 18,685 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 21.45s 24,501 2,369 1,084
Step 3.7 Flash 10.0 10.0 100.0% 0 7.98s 13,683 6,426 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.37s 10,533 324 0
Step 3.7 Flash 7.3 5.8 83.3% 1 2.29s 7,398 2,667 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 7.7 10.0 66.7% 0 1.17s 630 51 0
Step 3.7 Flash 5.3 7.2 44.4% 1 43.31s 828 104,487 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.87s 723 256 0
Step 3.7 Flash 3.4 9.3 0.0% 0 7.00s 525 4,604 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 1.57s 939 114 0
Step 3.7 Flash 9.8 10.0 100.0% 0 1.58s 735 1,857 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.43s 939 370 0
Step 3.7 Flash 5.5 9.9 33.3% 0 1.84s 756 3,564 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 4.17s 15,339 373 0
Step 3.7 Flash 10.0 10.0 100.0% 0 3.25s 7,746 1,360 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 3.0 10.0 0.0% 0 2.25s 273 24 0
Step 3.7 Flash 3.0 10.0 0.0% 0 124.75s 237 134,779 0

Quick Compare

Switch Comparison Pair