Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

ByteDance Seed: Seed-2.0-Lite vs Google: Gemini 3.1 Flash Lite

Last updated at: 2026-05-22

Metric Seed-2.0-Lite Seed-2.0-Lite medium Release: 2026-02-14 Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite high Release: 2026-05-08
Score 8.1 7.5
Rank #21 #48
Reliability 10.0 9.8
Consistency 8.9 8.1
Tests Correct
Attempt pass rate 75.0% 74.1%
Flaky tests 3 4
Total Runs 60 54
Cost per result 1.170 18.579
Total Cost $0.153 $2.044
Input Price $0.250 / 1M $0.250 / 1M
Output Price $2.000 / 1M $1.500 / 1M
Output Tokens 3,282 1,984
Reasoning Tokens 67,287 1,355,583
Response Time (avg) 36.79s 61.96s
Response Time (max) 168.71s 149.23s
Response Time (total) 735.86s 1115.31s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 8.3 10.0 75.0% 0 17.99s 996 7,142
Gemini 3.1 Flash Lite 9.4 10.0 100.0% 0 37.16s 100 130,598
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 7.0 9.7 50.0% 0 107.65s 452 20,524
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 137.63s 666 188,733
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 37.67s 506 4,299
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 149.23s 327 198,243
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 9.07s 246 1,742
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 4.49s 279 7,351
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 5.9 7.2 55.6% 1 88.74s 15 23,897
Gemini 3.1 Flash Lite 3.6 7.2 22.2% 1 139.90s 18 566,210
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 6.7 3.6 66.7% 1 18.25s 304 1,620
Gemini 3.1 Flash Lite 5.0 2.1 66.7% 1 45.69s 95 64,644
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 7.26s 71 1,480
Gemini 3.1 Flash Lite 7.3 5.8 83.3% 1 23.26s 52 3,549
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 9.0 7.9 88.9% 1 11.03s 461 3,532
Gemini 3.1 Flash Lite 5.7 6.8 44.4% 1 50.83s 213 193,654
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 10.0 10.0 100.0% 0 12.38s 222 1,011
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 6.44s 234 2,601
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Lite 3.0 10.0 0.0% 0 48.32s 9 2,040
Gemini 3.1 Flash Lite - - - - - - - -

Quick Compare

Switch Comparison Pair