Navigate
AI BENCHY
Advertise here

Mercury 2 vs Ling 3.0 Tiny

Mercury 2 leads on average score with 4.6 vs 4.0. Ling 3.0 Tiny has the lower benchmark cost at $0.000 vs $0.030. Mercury 2 is faster at 829ms vs 13.42s, with pass rates of 22.7% vs 9.1%.

Last updated at: 2026-08-07

Rank
#224
Total Output Tokens
9,564
Response Time (avg)
829ms
Total Cost
$0.030
Rank
#237
Total Output Tokens
110,077
Response Time (avg)
13.42s
Total Cost
$0.000
Recommended model Mercury 2

It has the best score here (4.6), while responding about 16.2x faster than Ling 3.0 Tiny.

Detailed comparison

Metric Mercury 2 Mercury 2 none Release: 2026-02-24 Ling 3.0 Tiny Ling 3.0 Tiny none Release: 2026-08-07 Free Available
Score 4.6 4.0
Rank #224 #237
Reliability 10.0 9.8
Consistency 9.2 10.0
Benchmark coverage 22/22 tests · 66/66 attempts 22/22 tests · 66/66 attempts
Tests Correct
Attempt pass rate 22.7% 9.1%
Flaky tests 2 0
Total Runs 66 66
Cost per result 0.734 0.000
Total Cost $0.030 $0.000
Input Price $0.250 / 1M $0.000 / 1M
Output Price $0.750 / 1M $0.000 / 1M
Total Input Tokens 88,704 184,762
Output Tokens 9,564 110,077
Reasoning Tokens 0 0
Response Time (avg) 829ms 13.42s
Response Time (max) 4.52s 240.61s
Response Time (total) 18.24s 295.29s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#224 Mercury 2

none
Cost
$0.002
Time
1.8s
Tokens
1,514 tok

#237 Ling 3.0 Tiny

none
Invalid SVG
Cost
$0.000
Time
334.1s
Tokens
32,913 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Mercury 2 3.4 9.6 0.0% 0 1.03s 7,229 3,088 0
Ling 3.0 Tiny 4.2 10.0 0.0% 0 1.19s 8,307 798 0

Quick Compare

Switch Comparison Pair