- Rank
- #253
- Total Output Tokens
- 943,044
- Response Time (avg)
- 74.52s
- Total Cost
- $0.645
Trinity Large Thinking (high) vs Space Bunny Alpha (low)
Trinity Large Thinking (high) leads on average score with 5.7 vs 5.7. Space Bunny Alpha (low) has the lower benchmark cost at $0.000 vs $0.645. Space Bunny Alpha (low) is faster at 30.63s vs 74.52s, with pass rates of 42.0% vs 53.6%.
Compared models
- Rank
- #259
- Total Output Tokens
- 194,450
- Response Time (avg)
- 30.63s
- Total Cost
- $0.000
Recommended model
Trinity Large Thinking (high)
It has the strongest score in this comparison (5.7) and the best overall balance of cost and response time across all 2 models.
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking high | Space Bunny Alpha Space Bunny Alpha low |
|---|---|---|
| Score | 5.7 | 5.7 |
| Rank | #253 | #259 |
| Reliability | 10.0 | 10.0 |
| Consistency | 7.3 | 6.5 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 42.0% | 53.6% |
| Flaky tests | 8 | 10 |
| Total Runs | 69 | 69 |
| Cost per result | 11.289 | 0.000 |
| Total Cost | $0.645 | $0.000 |
| Input Price | $0.250 / 1M | $0.000 / 1M |
| Output Price | $0.800 / 1M | $0.000 / 1M |
| Total Input Tokens | 283,116 | 289,086 |
| Output Tokens | 277,920 | 194,450 |
| Reasoning Tokens | 665,124 | 0 |
| Response Time (avg) | 74.52s | 30.63s |
| Response Time (max) | 510.21s | 340.88s |
| Response Time (total) | 1713.90s | 704.47s |
| Parameters | 398B total (13B active) | - |
| Availability | Weights available | Closed |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#253 Trinity Large Thinking
high- Cost
- $0.028
- Time
- 130.1s
- Tokens
- 34,387 tok
#259 Space Bunny Alpha
low- Cost
- $0.000
- Time
- 20.6s
- Tokens
- 3,462 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 44.97s | 181,340 | 1,553 | 8,872 | |
| Space Bunny Alpha | 4.7 | 3.1 | 33.3% | 1 | 89.00s | 169,129 | 10,896 | 0 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.7 | 8.1 | 58.3% | 1 | 4.13s | 663 | 3,108 | 7,565 | |
| Space Bunny Alpha | 6.9 | 5.8 | 75.0% | 2 | 2.75s | 2,130 | 2,255 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.7 | 4.7 | 33.3% | 2 | 245.04s | 7,204 | 83,616 | 266,836 | |
| Space Bunny Alpha | 5.6 | 4.7 | 55.6% | 2 | 6.60s | 8,493 | 5,897 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 262.39s | 76,626 | 14,706 | 149,243 | |
| Space Bunny Alpha | 3.8 | 5.8 | 33.3% | 1 | 12.17s | 85,984 | 8,071 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 5.8 | 66.7% | 1 | 13.98s | 6,906 | 580 | 9,338 | |
| Space Bunny Alpha | 10.0 | 10.0 | 100.0% | 0 | 1.35s | 7,890 | 401 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 7.2 | 11.1% | 1 | 66.94s | 756 | 163,364 | 166,091 | |
| Space Bunny Alpha | 5.3 | 7.2 | 44.4% | 1 | 65.69s | 1,869 | 61,188 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.5 | 3.4 | 66.7% | 1 | 12.90s | 501 | 5,700 | 5,942 | |
| Space Bunny Alpha | 4.4 | 9.9 | 0.0% | 0 | 2.78s | 855 | 427 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 10.0 | 50.0% | 0 | 4.12s | 684 | 88 | 2,853 | |
| Space Bunny Alpha | 4.4 | 2.7 | 33.3% | 2 | 1.43s | 1,425 | 351 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.7 | 5.2 | 33.3% | 2 | 3.84s | 678 | 2,566 | 3,368 | |
| Space Bunny Alpha | 7.9 | 9.9 | 66.7% | 0 | 3.51s | 1,782 | 1,939 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.96s | 7,551 | 282 | 1,440 | |
| Space Bunny Alpha | 4.7 | 1.6 | 66.7% | 1 | 3.48s | 8,953 | 370 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 125.13s | 207 | 2,357 | 43,576 | |
| Space Bunny Alpha | 3.0 | 10.0 | 0.0% | 0 | 340.88s | 576 | 102,655 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkinghighvsRing 2.6 1tmediumTrinity Large ThinkinghighvsGPT-5 NanomediumGPT-5.4nonevsSpace Bunny AlphalowTrinity Large ThinkinghighvsGPT-5.4 MininoneGemma 4 31BnoneFree AvailablevsSpace Bunny AlphalowTrinity Large ThinkinghighvsSeed 2.1 TurbononeTrinity Large ThinkinghighvsDots 3 Note PreviewmediumFree AvailableSpace Bunny AlphalowvsSolar Pro 4noneDots 3 Note PreviewmediumFree AvailablevsSpace Bunny AlphalowTrinity Large ThinkinghighvsDeepSeek V4.1 FlashnoneSeed 2.1 TurbononevsSpace Bunny AlphalowTrinity Large ThinkinghighvsGPT-5.6 Terranone