- Rank
- #232
- Total Output Tokens
- 1,016,785
- Response Time (avg)
- 85.44s
- Total Cost
- $0.756
Trinity Large Thinking (medium) vs Gemini 3.1 Flash Lite (minimal)
The average score is effectively tied at 6.1 vs 6.1. Gemini 3.1 Flash Lite (minimal) has the lower benchmark cost at $0.047 vs $0.756. Gemini 3.1 Flash Lite (minimal) is faster at 1.85s vs 85.44s, with pass rates of 48.5% vs 51.5%.
Compared models
- Rank
- #225
- Total Output Tokens
- 11,118
- Response Time (avg)
- 1.85s
- Total Cost
- $0.047
Recommended model
Gemini 3.1 Flash Lite (minimal)
It has the best score here (6.1), while costing about 16.3x less than Trinity Large Thinking (medium).
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking medium | Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite minimal |
|---|---|---|
| Score | 6.1 | 6.1 |
| Rank | #232 | #225 |
| Reliability | 10.0 | 10.0 |
| Consistency | 7.9 | 8.9 |
| Attempts | 66/66 | 66/66 |
| Tests Correct | ||
| Attempt pass rate | 48.5% | 51.5% |
| Flaky tests | 6 | 3 |
| Total Runs | 66 | 66 |
| Cost per result | 9.972 | 0.465 |
| Total Cost | $0.756 | $0.047 |
| Input Price | $0.250 / 1M | $0.250 / 1M |
| Output Price | $0.800 / 1M | $1.500 / 1M |
| Total Input Tokens | 118,486 | 119,076 |
| Output Tokens | 156,166 | 11,118 |
| Reasoning Tokens | 860,619 | 0 |
| Response Time (avg) | 85.44s | 1.85s |
| Response Time (max) | 525.14s | 12.97s |
| Response Time (total) | 1879.75s | 40.74s |
| Parameters | 398B total (13B active) | ~150B total (~10B active) |
| Availability | Weights available | Closed |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#232 Trinity Large Thinking
medium- Cost
- $0.016
- Time
- 75.9s
- Tokens
- 19,401 tok
#225 Gemini 3.1 Flash Lite
minimal- Cost
- $0.001
- Time
- 3.7s
- Tokens
- 635 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.8 | 9.8 | 50.0% | 0 | 5.11s | 663 | 1,531 | 6,078 | |
| Gemini 3.1 Flash Lite | 8.3 | 10.0 | 75.0% | 0 | 1.10s | 500 | 639 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.5 | 10.0 | 66.7% | 0 | 179.02s | 7,248 | 7,413 | 351,155 | |
| Gemini 3.1 Flash Lite | 5.5 | 10.0 | 33.3% | 0 | 831ms | 8,126 | 666 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 6.0 | 16.7% | 1 | 382.70s | 93,292 | 35,814 | 269,745 | |
| Gemini 3.1 Flash Lite | 3.0 | 10.0 | 0.0% | 0 | 7.75s | 94,962 | 8,988 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 5.8 | 66.7% | 1 | 16.38s | 6,906 | 475 | 11,153 | |
| Gemini 3.1 Flash Lite | 10.0 | 10.0 | 100.0% | 0 | 1.04s | 7,552 | 279 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.3 | 7.2 | 44.4% | 1 | 124.17s | 756 | 99,334 | 172,018 | |
| Gemini 3.1 Flash Lite | 2.9 | 7.2 | 11.1% | 1 | 972ms | 652 | 15 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 32.22s | 501 | 7,035 | 7,774 | |
| Gemini 3.1 Flash Lite | 4.0 | 10.0 | 0.0% | 0 | 791ms | 490 | 63 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.4 | 6.9 | 16.7% | 1 | 3.99s | 684 | 77 | 3,798 | |
| Gemini 3.1 Flash Lite | 10.0 | 10.0 | 100.0% | 0 | 932ms | 615 | 72 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.2 | 5.4 | 44.5% | 2 | 2.90s | 678 | 2,508 | 3,231 | |
| Gemini 3.1 Flash Lite | 6.0 | 4.6 | 66.7% | 2 | 2.15s | 564 | 153 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.26s | 7,551 | 299 | 1,372 | |
| Gemini 3.1 Flash Lite | 10.0 | 10.0 | 100.0% | 0 | 3.51s | 5,457 | 234 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 97.44s | 207 | 1,680 | 34,295 | |
| Gemini 3.1 Flash Lite | 3.0 | 10.0 | 0.0% | 0 | 724ms | 158 | 9 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkingmediumvsQwen3.6 FlashnoneTrinity Large ThinkingmediumvsGemma 4 31BnoneFree AvailableTrinity Large ThinkingmediumvsInklinglowClaude Sonnet 5nonevsGemini 3.1 Flash LiteminimalGemini 3.1 Flash LiteminimalvsQwen3.5-35B-A3BmediumTrinity Large ThinkingmediumvsGemini 3.1 Flash LitenoneTrinity Large ThinkingmediumvsQwen3.7 FlashnoneTrinity Large ThinkingmediumvsQwen3.5 Plus 2026-04-20noneGemini 3.1 Flash Liteminimalvsgpt-oss-120bmediumGemini 3.1 Flash LiteminimalvsQwen3.7 FlashnoneGemini 3.1 Flash LiteminimalvsGPT-5.6 LunalowGemini 3.1 Flash LiteminimalvsQwen3.5-Flashmedium