Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Hunter Alpha vs StepFun: Step 3.5 Flash

Last updated at: 2026-03-12

Metric Hunter Alpha Hunter Alpha none Release: Unknown release date Step 3.5 Flash Step 3.5 Flash medium Release: 2026-02-01 Free Available
Rank #50 #14
Avg Score 4.6 7.4
Consistency 8.0 9.1
Cost per result 0.000 0.000
Total Cost $0.000 $0.000
Tests Correct
Attempt pass rate 52.1% 68.8%
Flaky tests 4 2
Total Runs 48 48
Output Tokens 2,272 71,452
Reasoning Tokens 0 155,147
Response Time (avg) 4.64s 29.10s
Response Time (max) 15.17s 170.45s
Response Time (total) 74.24s 290.96s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Avg Score vs Response Time (avg)

Total Output Tokens

Avg Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 1.3 7.4 22.2% 1 3.85s 773 0
Step 3.5 Flash 10.0 10.0 100.0% 0 18.54s 13,924 17,208
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 10.0 10.0 0.0% 0 15.17s 379 0
Step 3.5 Flash 10.0 10.0 100.0% 0 29.57s 1,176 12,984
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 9.9 10.0 100.0% 0 8.49s 249 0
Step 3.5 Flash 10.0 10.0 100.0% 0 15.01s 600 13,886
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 4.0 10.0 33.3% 0 2.33s 27 0
Step 3.5 Flash 4.0 7.2 44.4% 1 170.45s 45,350 90,436
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 5.0 3.1 66.7% 1 2.71s 91 0
Step 3.5 Flash 6.0 10.0 0.0% 0 6.54s 2,214 2,584
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 5.0 10.0 50.0% 0 2.82s 69 0
Step 3.5 Flash 9.0 6.8 83.3% 1 4.98s 2,284 3,412
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 4.0 4.4 66.7% 2 3.06s 349 0
Step 3.5 Flash 4.0 10.0 33.3% 0 7.72s 5,629 10,835
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Hunter Alpha 10.0 10.0 100.0% 0 6.02s 335 0
Step 3.5 Flash 10.0 10.0 100.0% 0 11.91s 275 3,802

Quick Compare

Switch Comparison Pair