- Rank
- #222
- Total Output Tokens
- 842,200
- Response Time (avg)
- 95.16s
- Total Cost
- $0.700
Trinity Large Thinking (low) vs DeepSeek V4 Flash 0731
Trinity Large Thinking (low) leads on average score with 6.3 vs 6.2. DeepSeek V4 Flash 0731 has the lower benchmark cost at $0.055 vs $0.700. DeepSeek V4 Flash 0731 is faster at 27.42s vs 95.16s, with pass rates of 47.8% vs 30.4%.
Compared models
- Rank
- #228
- Total Output Tokens
- 39,789
- Response Time (avg)
- 27.42s
- Total Cost
- $0.055
Recommended model
DeepSeek V4 Flash 0731
Its score stays close to the best score here (6.2 vs 6.3), while costing about 12.9x less than Trinity Large Thinking (low).
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking low | DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 none |
|---|---|---|
| Score | 6.3 | 6.2 |
| Rank | #222 | #228 |
| Reliability | 10.0 | 10.0 |
| Consistency | 8.1 | 8.7 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 47.8% | 30.4% |
| Flaky tests | 6 | 4 |
| Total Runs | 69 | 69 |
| Cost per result | 9.163 | 0.826 |
| Total Cost | $0.700 | $0.055 |
| Input Price | $0.250 / 1M | $0.010 / 1M |
| Output Price | $0.800 / 1M | $1.280 / 1M |
| Total Input Tokens | 347,980 | 349,876 |
| Output Tokens | 121,898 | 39,789 |
| Reasoning Tokens | 720,302 | 0 |
| Response Time (avg) | 95.16s | 27.42s |
| Response Time (max) | 540.96s | 387.15s |
| Response Time (total) | 2188.74s | 630.70s |
| Parameters | 398B total (13B active) | 284B total (13B active) |
| Availability | Weights available | Closed |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#222 Trinity Large Thinking
low- Cost
- $0.021
- Time
- 173.2s
- Tokens
- 24,586 tok
#228 DeepSeek V4 Flash 0731
none- Cost
- $0.004
- Time
- 198.4s
- Tokens
- 14,316 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.4 | 3.8 | 66.7% | 1 | 63.23s | 225,126 | 2,088 | 21,753 | |
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 174.89s | 255,310 | 14,363 | 0 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 8.3 | 10.0 | 75.0% | 0 | 5.70s | 663 | 1,451 | 5,431 | |
| DeepSeek V4 Flash 0731 | 3.4 | 8.3 | 8.3% | 1 | 4.19s | 540 | 801 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.3 | 10.0 | 33.3% | 0 | 344.45s | 5,441 | 27,039 | 217,241 | |
| DeepSeek V4 Flash 0731 | 4.3 | 10.0 | 0.0% | 0 | 1.72s | 7,275 | 505 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.2 | 6.0 | 33.3% | 1 | 291.02s | 99,467 | 9,848 | 234,039 | |
| DeepSeek V4 Flash 0731 | 3.7 | 1.8 | 50.0% | 2 | 202.44s | 65,531 | 22,822 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.3 | 5.8 | 83.3% | 1 | 14.30s | 6,906 | 459 | 9,451 | |
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 2.37s | 7,290 | 199 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 7.2 | 11.1% | 1 | 140.26s | 756 | 72,049 | 217,598 | |
| DeepSeek V4 Flash 0731 | 5.3 | 10.0 | 33.3% | 0 | 1.17s | 675 | 26 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 12.65s | 501 | 5,138 | 5,695 | |
| DeepSeek V4 Flash 0731 | 4.4 | 9.9 | 0.0% | 0 | 3.22s | 471 | 153 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.4 | 6.8 | 16.7% | 1 | 4.07s | 684 | 45 | 3,710 | |
| DeepSeek V4 Flash 0731 | 5.0 | 6.8 | 33.3% | 1 | 1.42s | 627 | 69 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 7.9 | 44.4% | 1 | 3.28s | 678 | 2,625 | 3,077 | |
| DeepSeek V4 Flash 0731 | 3.2 | 9.9 | 0.0% | 0 | 2.74s | 594 | 302 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.19s | 7,551 | 299 | 1,430 | |
| DeepSeek V4 Flash 0731 | 9.5 | 10.0 | 100.0% | 0 | 4.61s | 11,380 | 536 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 2.10s | 207 | 857 | 877 | |
| DeepSeek V4 Flash 0731 | 3.0 | 10.0 | 0.0% | 0 | 1.85s | 183 | 13 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkinglowvsDeepSeek V3.2mediumTrinity Large ThinkinglowvsSpace Bunny AlphamediumDeepSeek V4 Flash 0731nonevsKAT-Coder-Pro V2.5mediumTrinity Large ThinkinglowvsDeepSeek V4 Flash 0423noneTrinity Large ThinkinglowvsGrok 4.20mediumDeepSeek V4 Flash 0731nonevsSpace Bunny AlphaxhighClaude Fable 5.1lowvsDeepSeek V4 Flash 0731noneTrinity Large ThinkinglowvsQwen3.5 Plus 2026-02-15noneTrinity Large ThinkinglowvsSolar Pro 4highTrinity Large ThinkinglowvsSpace Bunny AlphaxhighDeepSeek V4 Flash 0731nonevsGrok 4.20mediumTrinity Large ThinkinglowvsQwen3.6 27Bnone