- Rank
- #110
- Total Output Tokens
- 577,214
- Response Time (avg)
- 94.16s
- Total Cost
- $0.462
DeepSeek V4 Flash 0731 (medium) vs Grok 4.5 (low)
The average score is effectively tied at 7.8 vs 7.8. DeepSeek V4 Flash 0731 (medium) has the lower benchmark cost at $0.462 vs $1.345. Grok 4.5 (low) is faster at 18.14s vs 94.16s, with pass rates of 72.5% vs 73.9%.
Compared models
- Rank
- #109
- Total Output Tokens
- 121,415
- Response Time (avg)
- 18.14s
- Total Cost
- $1.345
Recommended model
Grok 4.5 (low)
It has the best score here (7.8), while responding about 5.2x faster than DeepSeek V4 Flash 0731 (medium).
Detailed comparison
| Metric | DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 medium | Grok 4.5 Grok 4.5 low |
|---|---|---|
| Score | 7.8 | 7.8 |
| Rank | #110 | #109 |
| Reliability | 10.0 | 10.0 |
| Consistency | 7.2 | 9.3 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 72.5% | 73.9% |
| Flaky tests | 8 | 2 |
| Total Runs | 69 | 69 |
| Cost per result | 1.176 | 8.404 |
| Total Cost | $0.462 | $1.345 |
| Input Price | $0.010 / 1M | $2.000 / 1M |
| Output Price | $1.280 / 1M | $6.000 / 1M |
| Total Input Tokens | 274,189 | 308,068 |
| Output Tokens | 66,689 | 9,218 |
| Reasoning Tokens | 510,525 | 112,197 |
| Response Time (avg) | 94.16s | 18.14s |
| Response Time (max) | 303.26s | 205.28s |
| Response Time (total) | 2165.72s | 417.21s |
| Parameters | 284B total (13B active) | ~1.7T total (~170B active) |
| Availability | Closed | Closed |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#110 DeepSeek V4 Flash 0731
medium- Cost
- $0.004
- Time
- 93.3s
- Tokens
- 13,252 tok
#109 SpaceXAI: Grok 4.5
low- Cost
- $0.011
- Time
- 12.4s
- Tokens
- 2,018 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 303.26s | 184,370 | 8,291 | 18,068 | |
| Grok 4.5 | 5.0 | 10.0 | 0.0% | 0 | 61.53s | 182,463 | 1,713 | 4,213 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 8.2 | 7.9 | 83.3% | 1 | 17.95s | 698 | 2,101 | 6,715 | |
| Grok 4.5 | 10.0 | 10.0 | 100.0% | 0 | 2.75s | 2,991 | 261 | 3,000 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 6.4 | 4.4 | 77.8% | 2 | 198.78s | 6,032 | 4,109 | 160,154 | |
| Grok 4.5 | 10.0 | 10.0 | 100.0% | 0 | 13.72s | 9,579 | 303 | 15,641 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 7.3 | 5.8 | 83.3% | 1 | 150.58s | 65,118 | 40,110 | 97,296 | |
| Grok 4.5 | 6.5 | 10.0 | 50.0% | 0 | 12.75s | 82,028 | 6,031 | 3,583 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 5.21s | 7,290 | 287 | 1,519 | |
| Grok 4.5 | 10.0 | 10.0 | 100.0% | 0 | 3.44s | 8,937 | 225 | 2,549 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 2.9 | 7.2 | 11.1% | 1 | 203.11s | 592 | 9,333 | 154,069 | |
| Grok 4.5 | 2.9 | 7.2 | 11.1% | 1 | 77.09s | 2,523 | 14 | 71,813 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 6.1 | 3.1 | 66.7% | 1 | 12.16s | 471 | 66 | 1,589 | |
| Grok 4.5 | 6.1 | 3.1 | 66.7% | 1 | 4.88s | 1,077 | 98 | 1,253 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 8.5 | 6.8 | 83.3% | 1 | 2.42s | 706 | 109 | 580 | |
| Grok 4.5 | 9.8 | 10.0 | 100.0% | 0 | 2.80s | 1,857 | 57 | 1,878 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 8.2 | 7.3 | 88.9% | 1 | 65.12s | 594 | 1,530 | 66,970 | |
| Grok 4.5 | 10.0 | 10.0 | 100.0% | 0 | 3.20s | 2,430 | 217 | 2,590 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 10.0 | 10.0 | 100.0% | 0 | 4.41s | 8,178 | 374 | 152 | |
| Grok 4.5 | 10.0 | 10.0 | 100.0% | 0 | 5.83s | 13,394 | 284 | 871 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 3.0 | 10.0 | 0.0% | 0 | 56.64s | 140 | 379 | 3,413 | |
| Grok 4.5 | 3.0 | 10.0 | 0.0% | 0 | 13.98s | 789 | 15 | 4,806 |
Quick Compare
Switch Comparison Pair
Qwen3.7 PlusnonevsGrok 4.5lowDeepSeek V4 Flash 0731mediumvsQwen3.7 PlusnoneStep 3.7 FlashmediumvsGrok 4.5lowGrok 4.5lowvsGLM 5.3highDeepSeek V4 Flash 0731mediumvsGLM 5.3highDeepSeek V4 Flash 0731mediumvsGPT-6 LunahighDeepSeek V4 Flash 0731mediumvsMuse Glimmer 30BhighGPT-6 LunahighvsGrok 4.5lowMuse Glimmer 30BhighvsGrok 4.5lowQwen3.5 Plus 2026-04-20mediumvsGrok 4.5lowGrok 4.5lowvsMiMo-V2.6-FlashmediumDeepSeek V4 Flash 0731mediumvsQwen3.7 Flashlow