Navigate
AI BENCHY
Advertise here

Ling 3.0 Tiny vs Grok 4.20

Grok 4.20 leads on average score with 4.1 vs 4.0. Ling 3.0 Tiny has the lower benchmark cost at $0.000 vs $0.057. Grok 4.20 is faster at 1.11s vs 13.42s, with pass rates of 9.1% vs 27.3%.

Last updated at: 2026-08-07

Rank
#237
Total Output Tokens
110,077
Response Time (avg)
13.42s
Total Cost
$0.000
Rank
#233
Total Output Tokens
1,923
Response Time (avg)
1.11s
Total Cost
$0.057
Recommended model Grok 4.20

It has the best score here (4.1), while responding about 12.1x faster than Ling 3.0 Tiny.

Detailed comparison

Metric Ling 3.0 Tiny Ling 3.0 Tiny none Release: 2026-08-07 Free Available Grok 4.20 Grok 4.20 none Release: 2026-03-31
Score 4.0 4.1
Rank #237 #233
Reliability 9.8 N/A
Consistency 10.0 8.1
Benchmark coverage 22/22 tests · 66/66 attempts 18/22 tests · 54/66 attempts
Tests Correct
Attempt pass rate 9.1% 27.3%
Flaky tests 0 0
Total Runs 66 54
Cost per result 0.000 1.570
Total Cost $0.000 $0.057
Input Price $0.000 / 1M $1.250 / 1M
Output Price $0.000 / 1M $2.500 / 1M
Total Input Tokens 184,762 41,313
Output Tokens 110,077 1,923
Reasoning Tokens 0 0
Response Time (avg) 13.42s 1.11s
Response Time (max) 240.61s 6.04s
Response Time (total) 295.29s 19.96s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#237 Ling 3.0 Tiny

none
Invalid SVG
Cost
$0.000
Time
334.1s
Tokens
32,913 tok

#233 xAI: Grok 4.20

none
Cost
$0.004
Time
6.5s
Tokens
1,367 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Ling 3.0 Tiny 4.2 10.0 0.0% 0 1.19s 8,307 798 0
Grok 4.20 1.1 3.1 0.0% 0 1.22s 1,074 312 0

Quick Compare

Switch Comparison Pair