Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

ByteDance Seed: Seed-2.0-Lite vs Qwen: Qwen3.6 Plus

The average score is effectively tied at 7.9 vs 7.8. Seed-2.0-Lite (medium) has the lower benchmark cost at $0.234 vs $0.405. Qwen3.6 Plus (medium) is faster at 43.12s vs 48.53s, with pass rates of 74.2% vs 71.2%.

Recommended modelSeed-2.0-Lite (medium)It has the best score here (7.9), while costing about 1.7x less than Qwen3.6 Plus (medium).

Last updated at: 2026-07-25

Metric Seed-2.0-Lite Seed-2.0-Lite medium Release: 2026-02-14 Qwen3.6 Plus Qwen3.6 Plus medium Release: 2026-04-20
Score 7.9 7.8
Rank #42 #44
Reliability 10.0 10.0
Consistency 8.6 9.3
Tests Correct
Attempt pass rate 74.2% 71.2%
Flaky tests 4 2
Total Runs 66 66
Cost per result 1.669 1.514
Total Cost $0.234 $0.405
Input Price $0.250 / 1M $0.325 / 1M
Output Price $2.000 / 1M $1.950 / 1M
Total Input Tokens 129,897 97,689
Output Tokens 12,533 6,412
Reasoning Tokens 88,047 184,825
Response Time (avg) 48.53s 43.12s
Response Time (max) 254.92s 291.55s
Response Time (total) 1067.74s 905.53s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#42 Seed-2.0-Lite

medium
Cost
$0.005
Time
86.7s
Tokens
2,354 tok

#44 Qwen3.6 Plus

medium
Cost
$0.024
Time
219.0s
Tokens
12,235 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 8.0 9.8 66.7% 0 156.74s 8,247 458 31,890
Qwen3.6 Plus 6.1 7.8 44.4% 1 153.12s 7,098 58 50,586

Quick Compare

Switch Comparison Pair