- Rank
- #222
- Total Output Tokens
- 842,200
- Response Time (avg)
- 95.16s
- Total Cost
- $0.700
Trinity Large Thinking (low) vs Inkling
Trinity Large Thinking (low) leads on average score with 6.3 vs 6.1. Inkling has the lower benchmark cost at $0.304 vs $0.700. Inkling is faster at 5.30s vs 95.16s, with pass rates of 47.8% vs 31.9%.
Compared models
- Rank
- #233
- Total Output Tokens
- 16,536
- Response Time (avg)
- 5.30s
- Total Cost
- $0.304
Recommended model
Inkling
Its score stays close to the best score here (6.1 vs 6.3), while costing about 2.3x less than Trinity Large Thinking (low).
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking low | Inkling Inkling none Free Available |
|---|---|---|
| Score | 6.3 | 6.1 |
| Rank | #222 | #233 |
| Reliability | 10.0 | 10.0 |
| Consistency | 8.1 | 9.6 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 47.8% | 31.9% |
| Flaky tests | 6 | 1 |
| Total Runs | 69 | 69 |
| Cost per result | 9.163 | 4.337 |
| Total Cost | $0.700 | $0.304 |
| Input Price | $0.250 / 1M | $1.000 / 1M |
| Output Price | $0.800 / 1M | $4.050 / 1M |
| Total Input Tokens | 347,980 | 236,599 |
| Output Tokens | 121,898 | 16,536 |
| Reasoning Tokens | 720,302 | 0 |
| Response Time (avg) | 95.16s | 5.30s |
| Response Time (max) | 540.96s | 48.02s |
| Response Time (total) | 2188.74s | 121.91s |
| Parameters | 398B total (13B active) | 975B total (41B active) |
| Availability | Weights available | Open source |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#222 Trinity Large Thinking
low- Cost
- $0.021
- Time
- 173.2s
- Tokens
- 24,586 tok
#233 Thinking Machines: Inkling
none
Provider returned error
- Cost
- $0.000
- Time
- 5.3s
- Tokens
- 0 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.4 | 3.8 | 66.7% | 1 | 63.23s | 225,126 | 2,088 | 21,753 | |
| Inkling | 10.0 | 10.0 | 100.0% | 0 | 45.64s | 132,479 | 5,982 | 0 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 8.3 | 10.0 | 75.0% | 0 | 5.70s | 663 | 1,451 | 5,431 | |
| Inkling | 4.8 | 10.0 | 25.0% | 0 | 1.43s | 678 | 261 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.3 | 10.0 | 33.3% | 0 | 344.45s | 5,441 | 27,039 | 217,241 | |
| Inkling | 4.5 | 10.0 | 0.0% | 0 | 1.01s | 7,356 | 436 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.2 | 6.0 | 33.3% | 1 | 291.02s | 99,467 | 9,848 | 234,039 | |
| Inkling | 2.9 | 5.8 | 16.7% | 1 | 25.74s | 80,195 | 8,777 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.3 | 5.8 | 83.3% | 1 | 14.30s | 6,906 | 459 | 9,451 | |
| Inkling | 10.0 | 10.0 | 100.0% | 0 | 1.14s | 7,056 | 279 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 7.2 | 11.1% | 1 | 140.26s | 756 | 72,049 | 217,598 | |
| Inkling | 5.3 | 10.0 | 33.3% | 0 | 1.18s | 786 | 45 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 12.65s | 501 | 5,138 | 5,695 | |
| Inkling | 5.0 | 10.0 | 0.0% | 0 | 859ms | 495 | 156 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.4 | 6.8 | 16.7% | 1 | 4.07s | 684 | 45 | 3,710 | |
| Inkling | 6.3 | 10.0 | 50.0% | 0 | 1.72s | 696 | 72 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 7.9 | 44.4% | 1 | 3.28s | 678 | 2,625 | 3,077 | |
| Inkling | 5.6 | 9.9 | 33.3% | 0 | 931ms | 696 | 291 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.19s | 7,551 | 299 | 1,430 | |
| Inkling | 3.0 | 10.0 | 0.0% | 0 | 2.50s | 5,949 | 213 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 2.10s | 207 | 857 | 877 | |
| Inkling | 3.0 | 10.0 | 0.0% | 0 | 670ms | 213 | 24 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkinglowvsDeepSeek V3.2mediumTrinity Large ThinkinglowvsSpace Bunny AlphamediumTrinity Large ThinkingmediumvsInklingnoneFree AvailableGranite 4.2 8BmediumvsInklingnoneFree AvailableTrinity Large ThinkinglowvsDeepSeek V4 Flash 0423noneTrinity Large ThinkinglowvsGrok 4.20mediumTrinity Large ThinkinglowvsQwen3.5 Plus 2026-02-15noneTrinity Large ThinkinglowvsSolar Pro 4highTrinity Large ThinkinglowvsSpace Bunny AlphaxhighTrinity Large ThinkinglowvsQwen3.6 27BnoneTrinity Large ThinkinglowvsKAT-Coder-Pro V2.5mediumTrinity Large ThinkinglowvsQwen3.5-35B-A3Bmedium