Navigate
AI BENCHY
Advertise here

Trinity Large Thinking — low vs medium

low leads on average score with 6.0 vs 6.0. low has the lower benchmark cost at $0.656 vs $0.792. medium is faster at 84.96s vs 96.78s, with pass rates of 48.5% vs 45.5%.

Last updated at: 2026-07-28

Rank
#145
Total Output Tokens
815,220
Response Time (avg)
96.78s
Total Cost
$0.656
Rank
#150
Total Output Tokens
1,002,646
Response Time (avg)
84.96s
Total Cost
$0.792
Recommended model low

It has the strongest score in this comparison (6.0) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric Trinity Large Thinking Trinity Large Thinking low Release: 2026-07-28 Trinity Large Thinking Trinity Large Thinking medium Release: 2026-07-28
Score 6.0 6.0
Rank #145 #150
Reliability 9.7 10.0
Consistency 7.9 7.6
Tests Correct
Attempt pass rate 48.5% 45.5%
Flaky tests 6 7
Total Runs 66 66
Cost per result 8.199 11.308
Total Cost $0.656 $0.792
Input Price $0.220 / 1M $0.220 / 1M
Output Price $0.850 / 1M $0.850 / 1M
Total Input Tokens 122,845 118,477
Output Tokens 117,704 148,823
Reasoning Tokens 697,516 853,823
Response Time (avg) 96.78s 84.96s
Response Time (max) 540.96s 525.14s
Response Time (total) 2129.13s 1869.16s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#145 Trinity Large Thinking

low
Cost
$0.021
Time
173.2s
Tokens
24,586 tok

#150 Trinity Large Thinking

medium
Cost
$0.016
Time
75.9s
Tokens
19,401 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Trinity Large Thinking 5.3 10.0 33.3% 0 344.45s 5,441 27,039 217,241
Trinity Large Thinking 7.5 10.0 66.7% 0 179.02s 7,248 7,413 351,155

Quick Compare

Switch Comparison Pair