Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

ByteDance Seed: Seed-2.0-Lite vs Qwen: Qwen3.7 Plus

Summary

Seed-2.0-Lite vs Qwen3.7 Plus benchmark comparison: Seed-2.0-Lite leads on average score with 8.5 vs 8.2. Seed-2.0-Lite has the lower benchmark cost at $0.175 vs $0.177. Qwen3.7 Plus is faster at 38.95s vs 47.07s, with pass rates of 76.2% vs 77.8%.

Recommended model: Qwen3.7 Plus - It offers the best overall trade-off: a competitive score (8.2), faster response than Seed-2.0-Lite, and balanced cost.

Last updated at: 2026-06-12

Metric Seed-2.0-Lite Seed-2.0-Lite medium Release: 2026-02-14 Qwen3.7 Plus Qwen3.7 Plus medium Release: 2026-06-03
Score 8.5 8.2
Rank #21 #28
Reliability 10.0 10.0
Consistency 9.0 9.1
Tests Correct
Attempt pass rate 76.2% 77.8%
Flaky tests 3 2
Total Runs 63 63
Cost per result 1.250 1.474
Total Cost $0.175 $0.177
Input Price $0.250 / 1M $0.320 / 1M
Output Price $2.000 / 1M $1.280 / 1M
Total Input Tokens 46,740 40,939
Output Tokens 3,230 2,125
Reasoning Tokens 78,406 125,754
Response Time (avg) 47.07s 38.95s
Response Time (max) 254.92s 178.04s
Response Time (total) 988.37s 817.85s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#21 Seed-2.0-Lite

medium
Cost
$0.005
Time
86.7s
Tokens
2,354 tok

#28 Qwen3.7 Plus

medium
Cost
$0.018
Time
193.2s
Tokens
10,821 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 8.3 10.0 75.0% 0 17.99s 942 996 7,142
Qwen3.7 Plus 10.0 10.0 100.0% 0 8.58s 672 195 5,065
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 8.0 9.8 66.7% 0 156.74s 8,247 458 31,890
Qwen3.7 Plus 6.1 6.6 55.6% 1 108.60s 6,472 414 43,576
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 37.67s 16,254 506 4,299
Qwen3.7 Plus 10.0 10.0 100.0% 0 65.24s 14,934 366 10,132
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 9.07s 8,562 246 1,742
Qwen3.7 Plus 10.0 10.0 100.0% 0 21.75s 7,782 270 6,713
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 5.9 7.2 55.6% 1 88.74s 843 15 23,897
Qwen3.7 Plus 3.6 7.2 22.2% 1 45.35s 771 57 27,073
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 6.7 3.6 66.7% 1 18.25s 582 304 1,620
Qwen3.7 Plus 10.0 10.0 100.0% 0 25.48s 516 123 3,998
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 7.26s 834 71 1,480
Qwen3.7 Plus 10.0 10.0 100.0% 0 16.13s 699 102 5,013
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 9.0 7.9 88.9% 1 10.23s 894 403 3,285
Qwen3.7 Plus 10.0 10.0 100.0% 0 16.38s 696 280 7,312
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 12.38s 9,306 222 1,011
Qwen3.7 Plus 10.0 10.0 100.0% 0 15.02s 8,193 292 1,831
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 3.0 10.0 0.0% 0 48.32s 276 9 2,040
Qwen3.7 Plus 3.0 10.0 0.0% 0 91.07s 204 26 15,041

Quick Compare

Switch Comparison Pair