Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

StepFun: Step 3.7 Flash vs Xiaomi: MiMo-V2.5

Last updated at: 2026-05-29

Metric Step 3.7 Flash Step 3.7 Flash low Release: 2026-05-29 MiMo-V2.5 MiMo-V2.5 medium Release: 2026-04-22
Score 7.4 7.4
Rank #60 #58
Reliability 10.0 10.0
Consistency 8.7 8.4
Tests Correct
Attempt pass rate 68.3% 70.0%
Flaky tests 3 4
Total Runs 60 60
Cost per result 2.796 2.876
Total Cost $0.336 $0.346
Input Price $0.200 / 1M $0.140 / 1M
Output Price $1.150 / 1M $0.280 / 1M
Output Tokens 285,209 2,806
Reasoning Tokens 0 161,888
Response Time (avg) 16.06s 20.35s
Response Time (max) 124.75s 97.49s
Response Time (total) 321.11s 406.94s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 8.7 7.9 91.7% 1 4.02s 10,896 0
MiMo-V2.5 10.0 10.0 100.0% 0 4.14s 281 1,739
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 10.0 10.0 100.0% 0 9.43s 14,569 0
MiMo-V2.5 6.9 6.2 66.7% 1 64.48s 536 44,967
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 10.0 10.0 100.0% 0 7.98s 6,426 0
MiMo-V2.5 10.0 10.0 100.0% 0 16.86s 363 7,609
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 7.3 5.8 83.3% 1 2.29s 2,667 0
MiMo-V2.5 2.7 5.7 16.7% 1 6.33s 306 5,714
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 5.3 7.2 44.4% 1 43.31s 104,487 0
MiMo-V2.5 5.3 10.0 33.3% 0 34.53s 507 49,478
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 3.4 9.3 0.0% 0 7.00s 4,604 0
MiMo-V2.5 5.4 2.5 66.7% 1 5.37s 121 418
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 9.8 10.0 100.0% 0 1.58s 1,857 0
MiMo-V2.5 9.9 10.0 100.0% 0 1.80s 88 801
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 5.5 9.9 33.3% 0 1.84s 3,564 0
MiMo-V2.5 8.2 7.2 88.9% 1 20.25s 279 33,254
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 10.0 10.0 100.0% 0 3.25s 1,360 0
MiMo-V2.5 10.0 10.0 100.0% 0 7.29s 303 2,424
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Step 3.7 Flash 3.0 10.0 0.0% 0 124.75s 134,779 0
MiMo-V2.5 3.0 10.0 0.0% 0 51.29s 22 15,484

Quick Compare

Switch Comparison Pair