Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Trinity Large Thinking (low) vs Qwen3.5-35B-A3B (medium)

Qwen3.5-35B-A3B (medium) leads on average score with 6.1 vs 6.0. Trinity Large Thinking (low) has the lower benchmark cost at $0.658 vs $1.140. Trinity Large Thinking (low) is faster at 96.61s vs 110.84s, with pass rates of 47.0% vs 65.2%.

Last updated at: 2026-08-20

Rank
#181
Total Output Tokens
818,359
Response Time (avg)
96.61s
Total Cost
$0.658
Rank
#170
Total Output Tokens
893,463
Response Time (avg)
110.84s
Total Cost
$1.140
Recommended model Trinity Large Thinking (low)

Its score stays close to the best score here (6.0 vs 6.1), while costing about 1.7x less than Qwen3.5-35B-A3B (medium).

Detailed comparison

Metric Trinity Large Thinking Trinity Large Thinking low Release: 2026-07-28 Qwen3.5-35B-A3B Qwen3.5-35B-A3B medium Release: 2026-02-24
Score 6.0 6.1
Rank #181 #170
Reliability 10.0 10.0
Consistency 8.3 7.6
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 47.0% 65.2%
Flaky tests 5 6
Total Runs 66 66
Cost per result 8.221 9.806
Total Cost $0.658 $1.140
Input Price $0.220 / 1M $0.250 / 1M
Output Price $0.850 / 1M $1.250 / 1M
Total Input Tokens 122,854 130,450
Output Tokens 119,810 54,857
Reasoning Tokens 698,549 838,606
Response Time (avg) 96.61s 110.84s
Response Time (max) 540.96s 950.25s
Response Time (total) 2125.51s 2438.58s
Parameters 398B total (13B active) 35B total (3B active)
Availability Weights available Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#181 Trinity Large Thinking

low
Cost
$0.021
Time
173.2s
Tokens
24,586 tok

#170 Qwen3.5-35B-A3B

medium
Cost
$0.009
Time
71.4s
Tokens
8,631 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Trinity Large Thinking 5.3 10.0 33.3% 0 344.45s 5,441 27,039 217,241
Qwen3.5-35B-A3B 5.9 9.3 33.3% 0 206.65s 4,106 23,844 111,462

Quick Compare

Switch Comparison Pair