Navigate
Advertise here

Trinity Large Thinking (low) vs Ling 3.1 Flash (low)

The average score is effectively tied at 6.3 vs 6.3. Ling 3.1 Flash (low) has the lower benchmark cost at $0.000 vs $0.918. Ling 3.1 Flash (low) is faster at 85.22s vs 95.16s, with pass rates of 47.8% vs 66.7%.

Last updated at: 2026-10-03

Compared models

Rank
#226
Total Output Tokens
842,200
Response Time (avg)
95.16s
Total Cost
$0.918
Rank
#222
Total Output Tokens
704,170
Response Time (avg)
85.22s
Total Cost
$0.000
Recommended model Trinity Large Thinking (low)

It has the strongest score in this comparison (6.3) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric Trinity Large Thinking Trinity Large Thinking low Release: 2026-07-28 Ling 3.1 Flash Ling 3.1 Flash low Release: 2026-10-03
Score 6.3 6.3
Rank #226 #222
Reliability 10.0 9.7
Consistency 8.1 7.8
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 47.8% 66.7%
Flaky tests 6 6
Total Runs 69 69
Cost per result 9.163 0.000
Total Cost $0.918 $0.000
Input Price $0.250 / 1M $0.000 / 1M
Output Price $0.800 / 1M $0.000 / 1M
Cache Read Price $0.060 / 1M N/A
Cache Write Price N/A N/A
Total Input Tokens 347,980 271,890
Output Tokens 121,898 190,841
Reasoning Tokens 720,302 513,329
Response Time (avg) 95.16s 85.22s
Response Time (max) 540.96s 432.81s
Response Time (total) 2188.74s 1960.12s
Parameters 398B total (13B active) 560B total (25B active)
Availability Weights available Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#226 Trinity Large Thinking

low
Cost
$0.021
Time
173.2s
Tokens
24,586 tok

#222 Ling 3.1 Flash

low
Cost
$0.000
Time
133.5s
Tokens
12,950 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Trinity Large Thinking 5.3 10.0 33.3% 0 344.45s 5,441 27,039 217,241
Ling 3.1 Flash 5.6 4.2 66.7% 2 107.30s 8,298 20,816 86,330

Quick Compare

Switch Comparison Pair