Navigate
AI BENCHY
Advertise here

Trinity Large Thinking (low) vs Step 3.5 Flash (medium)

The average score is effectively tied at 6.0 vs 6.0. Step 3.5 Flash (medium) has the lower benchmark cost at $0.108 vs $0.656. Trinity Large Thinking (low) is faster at 96.78s vs 174.22s, with pass rates of 48.5% vs 51.5%.

Last updated at: 2026-07-28

Rank
#145
Total Output Tokens
815,220
Response Time (avg)
96.78s
Total Cost
$0.656
Rank
#146
Total Output Tokens
402,554
Response Time (avg)
174.22s
Total Cost
$0.108
Recommended model Step 3.5 Flash (medium)

It has the best score here (6.0), while costing about 6.1x less than Trinity Large Thinking (low).

Detailed comparison

Metric Trinity Large Thinking Trinity Large Thinking low Release: 2026-07-28 Step 3.5 Flash Step 3.5 Flash medium Release: 2026-02-01
Score 6.0 6.0
Rank #145 #146
Reliability 9.7 9.2
Consistency 7.9 9.0
Tests Correct
Attempt pass rate 48.5% 51.5%
Flaky tests 6 1
Total Runs 66 63
Cost per result 8.199 0.540
Total Cost $0.656 $0.108
Input Price $0.220 / 1M $0.100 / 1M
Output Price $0.850 / 1M $0.300 / 1M
Total Input Tokens 122,845 65,707
Output Tokens 117,704 108,561
Reasoning Tokens 697,516 293,993
Response Time (avg) 96.78s 174.22s
Response Time (max) 540.96s 1597.85s
Response Time (total) 2129.13s 2613.32s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#145 Trinity Large Thinking

low
Cost
$0.021
Time
173.2s
Tokens
24,586 tok

#146 Step 3.5 Flash

medium
Cost
$0.008
Time
277.1s
Tokens
23,695 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Trinity Large Thinking 5.3 10.0 33.3% 0 344.45s 5,441 27,039 217,241
Step 3.5 Flash 2.4 5.2 0.0% 0 258.38s 2,211 13,207 22,429

Quick Compare

Switch Comparison Pair