Ling 3.0 Tiny (medium) leads on average score with 3.6 vs 1.9. Ling 3.0 Tiny (medium) has the lower benchmark cost at $0.000 vs $0.003. D1 is faster at 508ms vs 64.99s, with pass rates of 23.2% vs 11.6%.
Last updated at:
2026-10-02
Compared models
Rank
#360
Total Output Tokens
923,276
Response Time (avg)
64.99s
Total Cost
$0.000
Rank
#382
Total Output Tokens
0
Response Time (avg)
508ms
Total Cost
$0.003
Recommended modelLing 3.0 Tiny (medium)
It has the strongest score in this comparison (3.6) and the best overall balance of cost and response time across all 2 models.
D1D1none36/69 attempts. Choices supplied by adapters. Supports 12/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-10-02
D1D1none36/69 attempts. Choices supplied by adapters. Supports 12/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-10-02
Score
3.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#360
#382
Reliability
9.6First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
8.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
4.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
A test is fully passed only if every run passed for that test.Wrong answer: 10Response Time (avg)508msResponse Time (max)820msResponse Time (total)6.10sA test is fully passed only if every run passed for that test.…
Attempt pass rate
23.2%Attempt pass rate = passed attempts / total attempts across runs.…
11.6%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
3Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
69Total Runs…
36Total Runs…
Cost per result
0.000
0.112
Total Cost
$0.000
$0.003
Input Price
$0.000 / 1MInput Price…
$0.040 / 1MInput Price…
Output Price
$0.000 / 1MOutput Price…
$0.000 / 1MOutput Price…
Cache Read Price
N/ACache Read Price…
$0.040 / 1MCache Read Price…
Cache Write Price
N/ACache Write Price…
N/ACache Write Price…
Total Input Tokens
103,898Total Input Tokens…
55,956Total Input Tokens…
Output Tokens
148,304Output Tokens…
0Output Tokens…
Reasoning Tokens
774,972Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
64.99sResponse Time (avg)…
508msResponse Time (avg)…
Response Time (max)
262.20sResponse Time (max)…
820msResponse Time (max)…
Response Time (total)
1429.68sResponse Time (total)…
6.10sResponse Time (total)…
Parameters
7.9B total (1.3B active)
-
Availability
Open source
Closed
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#360 Ling 3.0 Tiny
medium
No output was saved. The original provider response or failure reason is unavailable.
Cost
$0.000
Time
177.4s
Tokens
6,873 tok
#382 LiquidAI: D1
none
No showcase result has been generated for this model yet.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)24msResponse Time (max)24msResponse Time (total)24msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Wrong answer: 1Response Time (avg)23.73sResponse Time (max)81.06sResponse Time (total)94.91sA test is fully passed only if every run passed for that test.…
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)431msResponse Time (max)505msResponse Time (total)1.29sA test is fully passed only if every run passed for that test.…
2.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 2Did not follow instructions: 1Response Time (avg)211.99sResponse Time (max)219.75sResponse Time (total)635.97sA test is fully passed only if every run passed for that test.…
2.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)456msResponse Time (max)577msResponse Time (total)911msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Wrong answer: 1Response Time (avg)139.78sResponse Time (max)262.20sResponse Time (total)279.57sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
2.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
16.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)4.41sResponse Time (max)6.08sResponse Time (total)8.82sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)418msResponse Time (max)518msResponse Time (total)836msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 2API error: 1Response Time (avg)81.40sResponse Time (max)82.11sResponse Time (total)162.80sA test is fully passed only if every run passed for that test.…
3.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)555msResponse Time (max)589msResponse Time (total)1.66sA test is fully passed only if every run passed for that test.…
3.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
2.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)31.77sResponse Time (max)31.77sResponse Time (total)31.77sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
6.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)3.00sResponse Time (max)3.81sResponse Time (total)5.99sA test is fully passed only if every run passed for that test.…
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)820msResponse Time (max)820msResponse Time (total)820msA test is fully passed only if every run passed for that test.…
3.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2No answer: 1Response Time (avg)34.65sResponse Time (max)88.47sResponse Time (total)103.95sA test is fully passed only if every run passed for that test.…
1.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)570msResponse Time (max)570msResponse Time (total)570msA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)4.91sResponse Time (max)4.91sResponse Time (total)4.91sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)100.96sResponse Time (max)100.96sResponse Time (total)100.96sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…