Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

DeepSeek V3.2 (medium) vs Ling 3.1 Flash (low)

The average score is effectively tied at 6.3 vs 6.3. Ling 3.1 Flash (low) has the lower benchmark cost at $0.000 vs $0.152. DeepSeek V3.2 (medium) is faster at 70.08s vs 85.22s, with pass rates of 62.3% vs 66.7%.

Last updated at: 2026-10-03

Compared models

Rank
#225
Total Output Tokens
143,160
Response Time (avg)
70.08s
Total Cost
$0.152
Rank
#222
Total Output Tokens
704,170
Response Time (avg)
85.22s
Total Cost
$0.000
Recommended model DeepSeek V3.2 (medium)

It has the strongest score in this comparison (6.3) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric DeepSeek V3.2 DeepSeek V3.2 medium Release: 2025-12-01 Ling 3.1 Flash Ling 3.1 Flash low Release: 2026-10-03
Score 6.3 6.3
Rank #225 #222
Reliability 10.0 9.7
Consistency 7.5 7.8
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 62.3% 66.7%
Flaky tests 7 6
Total Runs 69 69
Cost per result 1.250 0.000
Total Cost $0.152 $0.000
Input Price $0.280 / 1M $0.000 / 1M
Output Price $0.420 / 1M $0.000 / 1M
Cache Read Price $0.028 / 1M N/A
Cache Write Price N/A N/A
Total Input Tokens 306,662 271,890
Output Tokens 17,819 190,841
Reasoning Tokens 125,341 513,329
Response Time (avg) 70.08s 85.22s
Response Time (max) 376.10s 432.81s
Response Time (total) 1611.86s 1960.12s
Parameters 671B total (37B active) 560B total (25B active)
Availability Open source Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#225 DeepSeek V3.2

medium
Cost
$0.001
Time
53.6s
Tokens
1,932 tok

#222 Ling 3.1 Flash

low
Cost
$0.000
Time
133.5s
Tokens
12,950 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V3.2 6.0 7.2 55.6% 1 248.68s 5,717 649 52,014
Ling 3.1 Flash 5.6 4.2 66.7% 2 107.30s 8,298 20,816 86,330

Quick Compare

Switch Comparison Pair