Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

Mistral: Mistral Small 4 vs StepFun: Step 3.5 Flash

Last updated at: 2026-03-17

Metric Mistral Small 4 Mistral Small 4 none Release: 2026-03-16 Step 3.5 Flash Step 3.5 Flash medium Release: 2026-02-01 Free Available
Rank #61 #22
Score 5.3 7.9
Consistency 9.5 9.1
Cost per result 0.108 0.000
Total Cost $0.006 $0.000
Tests Correct
Attempt pass rate 33.3% 70.6%
Flaky tests 1 2
Total Runs 51 49
Output Tokens 1,624 71,904
Reasoning Tokens 0 155,607
Response Time (avg) 629ms 26.78s
Response Time (max) 1.72s 170.45s
Response Time (total) 10.70s 294.58s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 3.4 7.9 16.7% 1 395ms 182 0
Step 3.5 Flash 10.0 10.0 100.0% 0 13.56s 14,376 17,668
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 3.0 10.0 0.0% 0 1.72s 496 0
Step 3.5 Flash 10.0 10.0 100.0% 0 29.57s 1,176 12,984
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 10.0 10.0 100.0% 0 822ms 261 0
Step 3.5 Flash 10.0 10.0 100.0% 0 15.01s 600 13,886
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 5.3 10.0 33.3% 0 367ms 28 0
Step 3.5 Flash 5.3 7.2 44.4% 1 170.45s 45,350 90,436
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 4.0 10.0 0.0% 0 729ms 205 0
Step 3.5 Flash 5.5 10.0 0.0% 0 6.54s 2,214 2,584
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 6.5 10.0 50.0% 0 380ms 69 0
Step 3.5 Flash 8.5 6.8 83.3% 1 4.98s 2,284 3,412
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 3.1 9.9 0.0% 0 589ms 170 0
Step 3.5 Flash 5.3 10.0 33.3% 0 7.72s 5,629 10,835
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mistral Small 4 10.0 10.0 100.0% 0 1.40s 213 0
Step 3.5 Flash 10.0 10.0 100.0% 0 11.91s 275 3,802

Quick Compare

Switch Comparison Pair