Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Ling 3.1 Flash (low) vs Grok 4.20 (medium)

The average score is effectively tied at 6.3 vs 6.3. Ling 3.1 Flash (low) has the lower benchmark cost at $0.000 vs $1.246. Grok 4.20 (medium) is faster at 36.73s vs 85.22s, with pass rates of 66.7% vs 58.0%.

Last updated at: 2026-10-03

Compared models

Rank
#222
Total Output Tokens
704,170
Response Time (avg)
85.22s
Total Cost
$0.000
Rank
#228
Total Output Tokens
302,087
Response Time (avg)
36.73s
Total Cost
$1.246
Recommended model Grok 4.20 (medium)

It has the best score here (6.3), while responding about 2.3x faster than Ling 3.1 Flash (low).

Detailed comparison

Metric Ling 3.1 Flash Ling 3.1 Flash low Release: 2026-10-03 Grok 4.20 Grok 4.20 medium Release: 2026-03-31
Score 6.3 6.3
Rank #222 #228
Reliability 9.7 10.0
Consistency 7.8 8.2
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 66.7% 58.0%
Flaky tests 6 5
Total Runs 69 69
Cost per result 0.000 14.374
Total Cost $0.000 $1.246
Input Price $0.000 / 1M $1.250 / 1M
Output Price $0.000 / 1M $2.500 / 1M
Cache Read Price N/A $0.200 / 1M
Cache Write Price N/A N/A
Total Input Tokens 271,890 392,127
Output Tokens 190,841 7,754
Reasoning Tokens 513,329 294,333
Response Time (avg) 85.22s 36.73s
Response Time (max) 432.81s 199.66s
Response Time (total) 1960.12s 844.68s
Parameters 560B total (25B active) ~1T total (~100B active)
Availability Closed Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#222 Ling 3.1 Flash

low
Cost
$0.000
Time
133.5s
Tokens
12,950 tok

#228 SpaceXAI: Grok 4.20

medium
Cost
$0.041
Time
110.3s
Tokens
16,336 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Ling 3.1 Flash 5.6 4.2 66.7% 2 107.30s 8,298 20,816 86,330
Grok 4.20 6.3 6.6 55.6% 1 109.93s 8,307 268 103,150

Quick Compare

Switch Comparison Pair