Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

ByteDance Seed: Seed-2.0-Lite vs Qwen: Qwen3.5 Plus 2026-02-15

Last updated at: 2026-03-12

Metric Seed-2.0-Lite Seed-2.0-Lite medium Release: 2026-02-14 Qwen3.5 Plus 2026-02-15 Qwen3.5 Plus 2026-02-15 medium Release: 2026-02-15
Rank #3 #5
Avg Score 8.5 8.3
Consistency 8.7 9.5
Cost per result 0.870 1.264
Total Cost $0.105 $0.165
Tests Correct
Attempt pass rate 87.5% 85.4%
Flaky tests 3 1
Total Runs 48 48
Output Tokens 2,815 1,735
Reasoning Tokens 44,618 77,212
Response Time (avg) 29.39s 34.45s
Response Time (max) 168.71s 79.86s
Response Time (total) 470.29s 310.09s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Avg Score vs Response Time (avg)

Total Output Tokens

Avg Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 23.34s 990 7,037
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 10.37s 186 5,926
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 37.67s 506 4,299
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 46.85s 421 7,906
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 9.9 10.0 100.0% 0 9.07s 246 1,742
Qwen3.5 Plus 2026-02-15 9.9 10.0 100.0% 0 46.91s 270 14,916
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 4.0 7.2 55.6% 1 88.74s 15 23,897
Qwen3.5 Plus 2026-02-15 4.0 10.0 33.3% 0 17.50s 35 16,680
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 7.0 3.6 66.7% 1 18.25s 304 1,620
Qwen3.5 Plus 2026-02-15 10.0 1.6 66.7% 1 79.86s 73 8,675
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 7.26s 71 1,480
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 31.93s 101 7,704
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 9.3 7.9 88.9% 1 11.03s 461 3,532
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 34.57s 340 14,496
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 12.38s 222 1,011
Qwen3.5 Plus 2026-02-15 10.0 10.0 100.0% 0 7.54s 309 909

Quick Compare

Switch Comparison Pair