Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

Grok 4.6 (high) vs Seed 2.1 Turbo (low) vs Qwen3.8 2.4T A95B (low) vs Seed-2.0-Code (low) benchmark comparison: Grok 4.6 (high) leads on Score with 9.3. Grok 4.6 (high) leads on Reliability with 10.0. Seed-2.0-Code (low) has the lowest Total Cost at $0.767. Seed-2.0-Code (low) is fastest at 72.24s.

Last updated at: 2026-09-24

Compared models

Rank
#19
Total Output Tokens
282,217
Response Time (avg)
84.15s
Total Cost
$1.908
Rank
#28
Total Output Tokens
597,236
Response Time (avg)
163.90s
Total Cost
$1.546
Rank
#78
Total Output Tokens
495,541
Response Time (avg)
118.97s
Total Cost
$2.672
Rank
#106
Total Output Tokens
241,309
Response Time (avg)
72.24s
Total Cost
$0.767
Recommended model Grok 4.6 (high)

It has the strongest score in this comparison (9.3) and the best overall balance of cost and response time across all 4 models.

Detailed comparison

Metric Grok 4.6 Grok 4.6 high Release: 2026-08-12 Seed 2.1 Turbo Seed 2.1 Turbo low Release: 2026-08-12 Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B low Release: 2026-08-12 Seed-2.0-Code Seed-2.0-Code low Release: 2026-08-12
Score 9.3 9.1 8.1 7.7
Rank #19 #28 #78 #106
Reliability 10.0 10.0 9.4 8.5
Consistency 10.0 9.2 9.2 7.8
Attempts 66/66 66/66 66/66 66/66
Tests Correct
Attempt pass rate 86.4% 80.3% 72.7% 77.3%
Flaky tests 0 2 2 6
Total Runs 66 66 66 66
Cost per result 10.042 9.091 17.812 5.899
Total Cost $1.908 $1.546 $2.672 $0.767
Input Price $2.000 / 1M $0.500 / 1M $2.000 / 1M $0.500 / 1M
Output Price $6.000 / 1M $2.500 / 1M $6.000 / 1M $3.000 / 1M
Total Input Tokens 107,275 104,464 113,279 85,657
Output Tokens 5,094 6,844 121,045 8,142
Reasoning Tokens 277,123 590,392 374,496 233,167
Response Time (avg) 84.15s 163.90s 118.97s 72.24s
Response Time (max) 618.49s 739.98s 534.21s 485.92s
Response Time (total) 1851.26s 3605.87s 2617.41s 1589.31s
Parameters ~1.7T total (~170B active) ~200B total (~20B active) 2.4T total (95B active) ~200B total (~20B active)
Availability Closed Closed Weights available Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#19 SpaceXAI: Grok 4.6

high
Cost
$0.107
Time
265.1s
Tokens
18,001 tok

#28 Seed 2.1 Turbo

low
Cost
$0.049
Time
298.8s
Tokens
19,730 tok

#78 Qwen3.8 2.4T A95B

low
Cost
$0.100
Time
193.2s
Tokens
16,767 tok

#106 Seed-2.0-Code

low
Provider returned error
Cost
$0.000
Time
0.3s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Grok 4.6 10.0 10.0 100.0% 0 102.20s 9,579 379 51,829
Seed 2.1 Turbo 10.0 10.0 100.0% 0 360.85s 8,220 385 207,368
Qwen3.8 2.4T A95B 7.8 9.3 66.7% 0 172.53s 6,716 21,405 57,399
Seed-2.0-Code 7.0 7.1 55.6% 1 109.70s 7,948 455 55,759

Quick Compare

Switch Comparison Pair