- Rank
- #232
- Total Output Tokens
- 1,033,971
- Response Time (avg)
- 84.01s
- Total Cost
- $0.820
Trinity Large Thinking (medium) vs DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 leads on average score with 6.2 vs 6.1. DeepSeek V4 Flash 0731 has the lower benchmark cost at $0.055 vs $0.820. DeepSeek V4 Flash 0731 is faster at 27.42s vs 84.01s, with pass rates of 49.3% vs 30.4%.
Compared models
- Rank
- #228
- Total Output Tokens
- 39,789
- Response Time (avg)
- 27.42s
- Total Cost
- $0.055
Recommended model
DeepSeek V4 Flash 0731
It has the best score here (6.2), while costing about 15.1x less than Trinity Large Thinking (medium).
Detailed comparison
| Metric | Trinity Large Thinking Trinity Large Thinking medium | DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 none |
|---|---|---|
| Score | 6.1 | 6.2 |
| Rank | #232 | #228 |
| Reliability | 10.0 | 10.0 |
| Consistency | 7.7 | 8.7 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 49.3% | 30.4% |
| Flaky tests | 7 | 4 |
| Total Runs | 69 | 69 |
| Cost per result | 10.772 | 0.826 |
| Total Cost | $0.820 | $0.055 |
| Input Price | $0.250 / 1M | $0.010 / 1M |
| Output Price | $0.800 / 1M | $1.280 / 1M |
| Total Input Tokens | 319,352 | 349,876 |
| Output Tokens | 158,447 | 39,789 |
| Reasoning Tokens | 875,524 | 0 |
| Response Time (avg) | 84.01s | 27.42s |
| Response Time (max) | 525.14s | 387.15s |
| Response Time (total) | 1932.24s | 630.70s |
| Parameters | 398B total (13B active) | 284B total (13B active) |
| Availability | Weights available | Closed |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#232 Trinity Large Thinking
medium- Cost
- $0.016
- Time
- 75.9s
- Tokens
- 19,401 tok
#228 DeepSeek V4 Flash 0731
none- Cost
- $0.004
- Time
- 198.4s
- Tokens
- 14,316 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.1 | 3.1 | 66.7% | 1 | 52.50s | 200,866 | 2,281 | 14,905 | |
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 174.89s | 255,310 | 14,363 | 0 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.8 | 9.8 | 50.0% | 0 | 5.11s | 663 | 1,531 | 6,078 | |
| DeepSeek V4 Flash 0731 | 3.4 | 8.3 | 8.3% | 1 | 4.19s | 540 | 801 | 0 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 7.5 | 10.0 | 66.7% | 0 | 179.02s | 7,248 | 7,413 | 351,155 | |
| DeepSeek V4 Flash 0731 | 4.3 | 10.0 | 0.0% | 0 | 1.72s | 7,275 | 505 | 0 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 2.9 | 6.0 | 16.7% | 1 | 382.70s | 93,292 | 35,814 | 269,745 | |
| DeepSeek V4 Flash 0731 | 3.7 | 1.8 | 50.0% | 2 | 202.44s | 65,531 | 22,822 | 0 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 6.3 | 5.8 | 66.7% | 1 | 16.38s | 6,906 | 475 | 11,153 | |
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 2.37s | 7,290 | 199 | 0 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.3 | 7.2 | 44.4% | 1 | 124.17s | 756 | 99,334 | 172,018 | |
| DeepSeek V4 Flash 0731 | 5.3 | 10.0 | 33.3% | 0 | 1.17s | 675 | 26 | 0 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 32.22s | 501 | 7,035 | 7,774 | |
| DeepSeek V4 Flash 0731 | 4.4 | 9.9 | 0.0% | 0 | 3.22s | 471 | 153 | 0 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 4.4 | 6.9 | 16.7% | 1 | 3.99s | 684 | 77 | 3,798 | |
| DeepSeek V4 Flash 0731 | 5.0 | 6.8 | 33.3% | 1 | 1.42s | 627 | 69 | 0 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 5.2 | 5.4 | 44.5% | 2 | 2.90s | 678 | 2,508 | 3,231 | |
| DeepSeek V4 Flash 0731 | 3.2 | 9.9 | 0.0% | 0 | 2.74s | 594 | 302 | 0 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 10.0 | 10.0 | 100.0% | 0 | 5.26s | 7,551 | 299 | 1,372 | |
| DeepSeek V4 Flash 0731 | 9.5 | 10.0 | 100.0% | 0 | 4.61s | 11,380 | 536 | 0 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | 3.0 | 10.0 | 0.0% | 0 | 97.44s | 207 | 1,680 | 34,295 | |
| DeepSeek V4 Flash 0731 | 3.0 | 10.0 | 0.0% | 0 | 1.85s | 183 | 13 | 0 |
Quick Compare
Switch Comparison Pair
Trinity Large ThinkingmediumvsInklingnoneFree AvailableTrinity Large ThinkingmediumvsQwen3.5 Plus 2026-04-20noneDeepSeek V4 Flash 0731nonevsKAT-Coder-Pro V2.5mediumTrinity Large ThinkingmediumvsGemini 3.5 FlashnoneDeepSeek V4 Flash 0731nonevsSpace Bunny AlphaxhighClaude Fable 5.1lowvsDeepSeek V4 Flash 0731noneTrinity Large ThinkingmediumvsKAT-Coder-Pro V2.5noneDeepSeek V4 Flash 0731nonevsGrok 4.20mediumClaude Fable 5.1lowvsTrinity Large ThinkingmediumTrinity Large ThinkinglowvsDeepSeek V4 Flash 0731noneTrinity Large ThinkingmediumvsGemini 3.1 Flash Lite PreviewnoneDeepSeek V4 Flash 0731nonevsSpace Bunny Alphamedium