Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

ByteDance Seed: Seed-2.0-Lite vs Qwen: Qwen3.7 Plus

Summary

Seed-2.0-Lite vs Qwen3.7 Plus benchmark comparison: Qwen3.7 Plus leads on average score with 6.4 vs 5.8. Seed-2.0-Lite has the lower benchmark cost at $0.019 vs $0.028. Seed-2.0-Lite is faster at 2.49s vs 2.85s, with pass rates of 46.0% vs 47.6%.

Recommended model: Seed-2.0-Lite - Its score stays close to the best score here (5.8 vs 6.4), while costing about 1.5x less than Qwen3.7 Plus.

Last updated at: 2026-06-10

Metric Seed-2.0-Lite Seed-2.0-Lite none Release: 2026-02-14 Qwen3.7 Plus Qwen3.7 Plus none Release: 2026-06-03
Score 5.8 6.4
Rank #111 #89
Reliability 10.0 10.0
Consistency 8.4 10.0
Tests Correct
Attempt pass rate 46.0% 47.6%
Flaky tests 4 0
Total Runs 63 63
Cost per result 0.228 0.276
Total Cost $0.019 $0.028
Input Price $0.250 / 1M $0.400 / 1M
Output Price $2.000 / 1M $1.600 / 1M
Total Input Tokens 46,573 42,510
Output Tokens 3,259 6,578
Reasoning Tokens 0 0
Response Time (avg) 2.49s 2.85s
Response Time (max) 6.70s 29.38s
Response Time (total) 52.26s 59.86s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#111 Seed-2.0-Lite

none
Cost
$0.005
Time
83.8s
Tokens
2,311 tok

#89 Qwen3.7 Plus

none
Cost
$0.019
Time
213.5s
Tokens
11,960 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 3.0 5.9 16.7% 2 2.43s 894 709 0
Qwen3.7 Plus 6.5 10.0 50.0% 0 1.38s 696 349 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 5.6 10.0 33.3% 0 2.83s 8,215 410 0
Qwen3.7 Plus 5.5 10.0 33.3% 0 2.15s 7,911 639 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 3.0 10.0 0.0% 0 6.59s 16,215 498 0
Qwen3.7 Plus 10.0 10.0 100.0% 0 29.38s 14,952 4,505 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 1.82s 8,538 246 0
Qwen3.7 Plus 10.0 10.0 100.0% 0 1.43s 7,794 243 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 3.6 7.2 22.2% 1 1.33s 939 17 0
Qwen3.7 Plus 3.0 10.0 0.0% 0 868ms 789 18 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 3.45s 570 294 0
Qwen3.7 Plus 5.3 10.0 0.0% 0 1.33s 522 78 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 1.06s 810 73 0
Qwen3.7 Plus 6.3 10.0 50.0% 0 929ms 711 72 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 5.3 7.2 44.4% 1 2.78s 858 709 0
Qwen3.7 Plus 7.7 10.0 66.7% 0 1.71s 714 443 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 3.94s 9,270 292 0
Qwen3.7 Plus 10.0 10.0 100.0% 0 3.54s 8,211 222 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Seed-2.0-Lite 3.0 10.0 0.0% 0 1.96s 264 11 0
Qwen3.7 Plus 3.0 10.0 0.0% 0 1.21s 210 9 0

Quick Compare

Switch Comparison Pair