Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Sonnet 4.6 vs StepFun: Step 3.7 Flash

Last updated at: 2026-05-29

Metric Claude Sonnet 4.6 Claude Sonnet 4.6 none Release: 2026-02-17 Step 3.7 Flash Step 3.7 Flash high Release: 2026-05-29
Score 7.0 7.1
Rank #78 #74
Reliability 10.0 10.0
Consistency 9.7 8.2
Tests Correct
Attempt pass rate 58.3% 65.8%
Flaky tests 1 4
Total Runs 60 60
Cost per result 2.782 8.723
Total Cost $0.306 $0.960
Input Price $3.000 / 1M $0.200 / 1M
Output Price $15.000 / 1M $1.150 / 1M
Output Tokens 9,450 828,084
Reasoning Tokens 0 0
Response Time (avg) 5.27s 49.43s
Response Time (max) 23.84s 192.75s
Response Time (total) 68.50s 988.58s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 4.8 10.0 25.0% 0 2.94s 1,214 0
Step 3.7 Flash 10.0 10.0 100.0% 0 13.40s 42,656 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.8 10.0 50.0% 0 6.73s 2,112 0
Step 3.7 Flash 3.6 4.6 25.0% 1 126.82s 164,069 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 9.5 10.0 100.0% 0 23.84s 3,766 0
Step 3.7 Flash 10.0 10.0 100.0% 0 13.01s 8,802 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 3.43s 252 0
Step 3.7 Flash 10.0 10.0 100.0% 0 14.72s 23,113 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 7.7 10.0 66.7% 0 3.54s 413 0
Step 3.7 Flash 4.1 4.4 44.5% 2 149.64s 410,502 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.1 3.1 66.7% 1 2.56s 192 0
Step 3.7 Flash 5.5 10.0 0.0% 0 4.17s 2,862 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.5 10.0 50.0% 0 1.96s 90 0
Step 3.7 Flash 9.8 10.0 100.0% 0 1.52s 2,010 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 7.7 10.0 66.7% 0 2.53s 533 0
Step 3.7 Flash 5.3 7.2 44.4% 1 10.22s 25,422 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 4.11s 447 0
Step 3.7 Flash 10.0 10.0 100.0% 0 2.79s 1,172 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 3.0 10.0 0.0% 0 4.67s 431 0
Step 3.7 Flash 3.0 10.0 0.0% 0 149.34s 147,476 0

Quick Compare

Switch Comparison Pair