- Rank
- #232
- Total Output Tokens
- 1,033,971
- Response Time (avg)
- 84.01s
- Total Cost
- $0.820
Trinity Large Thinking (medium) vs Inkling
The average score is effectively tied at 6.1 vs 6.1. Inkling has the lower benchmark cost at $0.304 vs $0.820. Inkling is faster at 5.30s vs 84.01s, with pass rates of 49.3% vs 31.9%.
Compared models
- Rank
- #233
- Total Output Tokens
- 16,536
- Response Time (avg)
- 5.30s
- Total Cost
- $0.304
Recommended model
Inkling
It has the best score here (6.1), while costing about 2.7x less than Trinity Large Thinking (medium).
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking medium | Inkling Inkling none Free Available |
|---|---|---|
| Score | 6.1 | 6.1 |
| Rank | #232 | #233 |
| Reliability | 10.0 | 10.0 |
| Consistency | 7.7 | 9.6 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 49.3% | 31.9% |
| Flaky tests | 7 | 1 |
| Total Runs | 69 | 69 |
| Cost per result | 10.772 | 4.337 |
| Total Cost | $0.820 | $0.304 |
| Input Price | $0.250 / 1M | $1.000 / 1M |
| Output Price | $0.800 / 1M | $4.050 / 1M |
| Total Input Tokens | 319,352 | 236,599 |
| Output Tokens | 158,447 | 16,536 |
| Reasoning Tokens | 875,524 | 0 |
| Response Time (avg) | 84.01s | 5.30s |
| Response Time (max) | 525.14s | 48.02s |
| Response Time (total) | 1932.24s | 121.91s |
| Parameters | 398B total (13B active) | 975B total (41B active) |
| Availability | Weights available | Open source |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#232 Trinity Large Thinking
medium- Cost
- $0.016
- Time
- 75.9s
- Tokens
- 19,401 tok
#233 Thinking Machines: Inkling
none
Provider returned error
- Cost
- $0.000
- Time
- 5.3s
- Tokens
- 0 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.1 | 3.1 | 66.7% | 1 | 52.50s | 200,866 | 2,281 | 14,905 | |
| Inkling | 10.0 | 10.0 | 100.0% | 0 | 45.64s | 132,479 | 5,982 | 0 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.8 | 9.8 | 50.0% | 0 | 5.11s | 663 | 1,531 | 6,078 | |
| Inkling | 4.8 | 10.0 | 25.0% | 0 | 1.43s | 678 | 261 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.5 | 10.0 | 66.7% | 0 | 179.02s | 7,248 | 7,413 | 351,155 | |
| Inkling | 4.5 | 10.0 | 0.0% | 0 | 1.01s | 7,356 | 436 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 6.0 | 16.7% | 1 | 382.70s | 93,292 | 35,814 | 269,745 | |
| Inkling | 2.9 | 5.8 | 16.7% | 1 | 25.74s | 80,195 | 8,777 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 5.8 | 66.7% | 1 | 16.38s | 6,906 | 475 | 11,153 | |
| Inkling | 10.0 | 10.0 | 100.0% | 0 | 1.14s | 7,056 | 279 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.3 | 7.2 | 44.4% | 1 | 124.17s | 756 | 99,334 | 172,018 | |
| Inkling | 5.3 | 10.0 | 33.3% | 0 | 1.18s | 786 | 45 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 32.22s | 501 | 7,035 | 7,774 | |
| Inkling | 5.0 | 10.0 | 0.0% | 0 | 859ms | 495 | 156 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.4 | 6.9 | 16.7% | 1 | 3.99s | 684 | 77 | 3,798 | |
| Inkling | 6.3 | 10.0 | 50.0% | 0 | 1.72s | 696 | 72 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.2 | 5.4 | 44.5% | 2 | 2.90s | 678 | 2,508 | 3,231 | |
| Inkling | 5.6 | 9.9 | 33.3% | 0 | 931ms | 696 | 291 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.26s | 7,551 | 299 | 1,372 | |
| Inkling | 3.0 | 10.0 | 0.0% | 0 | 2.50s | 5,949 | 213 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 97.44s | 207 | 1,680 | 34,295 | |
| Inkling | 3.0 | 10.0 | 0.0% | 0 | 670ms | 213 | 24 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkingmediumvsQwen3.5 Plus 2026-04-20noneGranite 4.2 8BmediumvsInklingnoneFree AvailableTrinity Large ThinkingmediumvsGemini 3.5 FlashnoneTrinity Large ThinkingmediumvsKAT-Coder-Pro V2.5noneClaude Fable 5.1lowvsTrinity Large ThinkingmediumClaude Fable 5.1lowvsInklingnoneFree AvailableTrinity Large ThinkingmediumvsGemini 3.1 Flash Lite PreviewnoneTrinity Large ThinkingmediumvsDeepSeek V4 Flash 0731noneInklingnoneFree AvailablevsGLM 5.3 FlashXlowTrinity Large ThinkingmediumvsGLM 5.3 FlashXlowTrinity Large ThinkingmediumvsSpace Bunny AlphaxhighTrinity Large ThinkingmediumvsMiMo-V2.6-Pronone