Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

Grok 4.6 (high) vs Seed 2.1 Turbo (low) vs Qwen3.8 2.4T A95B (low) vs Seed-2.0-Code (low) benchmark comparison: Grok 4.6 (high) leads on Score with 9.2. Grok 4.6 (high) leads on Reliability with 10.0. Seed-2.0-Code (low) has the lowest Total Cost at $0.681. Seed-2.0-Code (low) is fastest at 65.89s.

Last updated at: 2026-08-12

Rank
#11
Total Output Tokens
265,604
Response Time (avg)
79.38s
Total Cost
$1.809
Rank
#15
Total Output Tokens
631,257
Response Time (avg)
168.41s
Total Cost
$1.631
Rank
#46
Total Output Tokens
554,485
Response Time (avg)
132.39s
Total Cost
$3.113
Rank
#66
Total Output Tokens
212,558
Response Time (avg)
65.89s
Total Cost
$0.681
Recommended model Grok 4.6 (high)

It has the best score here (9.2), while responding about 1.5x faster than the other models in this comparison.

Detailed comparison

Metric Grok 4.6 Grok 4.6 high Release: 2026-08-12 Seed 2.1 Turbo Seed 2.1 Turbo low Release: 2026-08-12 Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B low Release: 2026-08-12 Seed-2.0-Code Seed-2.0-Code low Release: 2026-08-12
Score 9.2 9.1 8.2 7.7
Rank #11 #15 #46 #66
Reliability 10.0 9.9 9.1 5.7
Consistency 9.6 8.9 8.9 7.5
Attempts 66/66 66/66 66/66 66/66
Tests Correct
Attempt pass rate 84.9% 83.3% 78.8% 80.3%
Flaky tests 1 3 3 7
Total Runs 66 66 66 66
Cost per result 10.046 9.591 19.453 5.235
Total Cost $1.809 $1.631 $3.113 $0.681
Input Price $2.000 / 1M $0.500 / 1M $2.000 / 1M $0.500 / 1M
Output Price $6.000 / 1M $2.500 / 1M $6.000 / 1M $3.000 / 1M
Total Input Tokens 107,266 104,483 113,340 85,648
Output Tokens 5,094 6,848 107,311 8,142
Reasoning Tokens 260,510 624,409 447,174 204,416
Response Time (avg) 79.38s 168.41s 132.39s 65.89s
Response Time (max) 618.49s 739.98s 534.21s 485.92s
Response Time (total) 1746.29s 3705.09s 2912.56s 1449.53s
Parameters ~1.7T total (~170B active) ~200B total (~20B active) 2.4T total (95B active) ~200B total (~20B active)
Availability Closed Closed Weights available Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#11 SpaceXAI: Grok 4.6

high
Cost
$0.107
Time
265.1s
Tokens
18,001 tok

#15 Seed 2.1 Turbo

low
Cost
$0.049
Time
298.8s
Tokens
19,730 tok

#46 Qwen3.8 2.4T A95B

low
Cost
$0.100
Time
193.2s
Tokens
16,767 tok

#66 Seed-2.0-Code

low
Provider returned error
Cost
$0.000
Time
0.3s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Grok 4.6 10.0 10.0 100.0% 0 102.20s 9,579 379 51,829
Seed 2.1 Turbo 10.0 10.0 100.0% 0 360.85s 8,220 385 207,368
Qwen3.8 2.4T A95B 7.6 7.2 77.8% 1 172.53s 6,716 21,405 57,399
Seed-2.0-Code 7.0 5.0 77.8% 2 109.70s 7,948 455 55,759

Quick Compare

Switch Comparison Pair