- Rank
- #288
- Total Output Tokens
- 932,619
- Response Time (avg)
- 75.86s
- Total Cost
- $0.592
Trinity Large Thinking (high) vs Qwen3 Coder Next (medium)
Trinity Large Thinking (high) leads on average score with 4.8 vs 4.6. Qwen3 Coder Next (medium) has the lower benchmark cost at $0.034 vs $0.592. Qwen3 Coder Next (medium) is faster at 9.07s vs 75.86s, with pass rates of 39.4% vs 25.8%.
Compared models
- Rank
- #299
- Total Output Tokens
- 19,069
- Response Time (avg)
- 9.07s
- Total Cost
- $0.034
Recommended model
Qwen3 Coder Next (medium)
Its score stays close to the best score here (4.6 vs 4.8), while costing about 17.9x less than Trinity Large Thinking (high).
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking high | Qwen3 Coder Next Qwen3 Coder Next medium |
|---|---|---|
| Score | 4.8 | 4.6 |
| Rank | #288 | #299 |
| Reliability | 10.0 | 10.0 |
| Consistency | 7.2 | 8.6 |
| Attempts | 66/66 | 66/66 |
| Tests Correct | ||
| Attempt pass rate | 39.4% | 25.8% |
| Flaky tests | 8 | 4 |
| Total Runs | 66 | 66 |
| Cost per result | 12.473 | 1.058 |
| Total Cost | $0.592 | $0.034 |
| Input Price | $0.250 / 1M | $0.120 / 1M |
| Output Price | $0.800 / 1M | $0.800 / 1M |
| Total Input Tokens | 101,776 | 148,203 |
| Output Tokens | 276,367 | 19,069 |
| Reasoning Tokens | 656,252 | 0 |
| Response Time (avg) | 75.86s | 9.07s |
| Response Time (max) | 510.21s | 81.80s |
| Response Time (total) | 1668.93s | 154.19s |
| Parameters | 398B total (13B active) | 80B total (3B active) |
| Availability | Weights available | Open source |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#288 Trinity Large Thinking
high- Cost
- $0.028
- Time
- 130.1s
- Tokens
- 34,387 tok
#299 Qwen3 Coder Next
medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
- Cost
- $0.000
- Time
- 300.0s
- Tokens
- 0 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.7 | 8.1 | 58.3% | 1 | 4.13s | 663 | 3,108 | 7,565 | |
| Qwen3 Coder Next | 3.5 | 8.1 | 16.7% | 1 | 8.64s | 645 | 1,252 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.7 | 4.7 | 33.3% | 2 | 245.04s | 7,204 | 83,616 | 266,836 | |
| Qwen3 Coder Next | 3.7 | 7.2 | 22.2% | 1 | 924ms | 7,185 | 336 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 262.39s | 76,626 | 14,706 | 149,243 | |
| Qwen3 Coder Next | 3.0 | 10.0 | 0.0% | 0 | 14.65s | 121,413 | 16,067 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 5.8 | 66.7% | 1 | 13.98s | 6,906 | 580 | 9,338 | |
| Qwen3 Coder Next | 6.5 | 10.0 | 50.0% | 0 | 81.80s | 7,758 | 246 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 7.2 | 11.1% | 1 | 66.94s | 756 | 163,364 | 166,091 | |
| Qwen3 Coder Next | 3.6 | 7.2 | 22.2% | 1 | 570ms | 762 | 25 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.5 | 3.4 | 66.7% | 1 | 12.90s | 501 | 5,700 | 5,942 | |
| Qwen3 Coder Next | 6.3 | 3.4 | 66.7% | 1 | 1.39s | 498 | 142 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 10.0 | 50.0% | 0 | 4.12s | 684 | 88 | 2,853 | |
| Qwen3 Coder Next | 6.3 | 10.0 | 50.0% | 0 | 7.49s | 684 | 63 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.7 | 5.2 | 33.3% | 2 | 3.84s | 678 | 2,566 | 3,368 | |
| Qwen3 Coder Next | 3.0 | 10.0 | 0.0% | 0 | 1.25s | 678 | 671 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.96s | 7,551 | 282 | 1,440 | |
| Qwen3 Coder Next | 10.0 | 10.0 | 100.0% | 0 | 2.64s | 8,364 | 255 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 125.13s | 207 | 2,357 | 43,576 | |
| Qwen3 Coder Next | 3.0 | 10.0 | 0.0% | 0 | 399ms | 216 | 12 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkinghighvsGPT-5.4 NanononeTrinity Large ThinkinghighvsNemotron 3.5 LightninglowFree AvailableGranite 4.2 8BhighvsQwen3 Coder NextmediumTrinity Large ThinkinghighvsRing-2.6-1TnoneTrinity Large ThinkinghighvsKAT-Coder-Air V2.5noneTrinity Large ThinkinghighvsHy4 previewnoneTrinity Large ThinkinghighvsGLM 4.7 FlashnoneTrinity Large ThinkinghighvsNemotron 3 SupernoneFree AvailableLaguna S 2.1noneFree AvailablevsQwen3 Coder NextmediumTrinity Large ThinkinghighvsLing-2.6-flashnoneMercury 2nonevsQwen3 Coder NextmediumTrinity Large ThinkinghighvsCobuddymedium