Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

ByteDance Seed: Seed-2.0-Mini vs Google: Gemini 3.1 Flash Lite

Last updated at: 2026-05-29

Metric Seed-2.0-Mini Seed-2.0-Mini medium Release: 2026-02-14 Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite low Release: 2026-05-08
Score 7.1 7.4
Rank #75 #55
Reliability 10.0 10.0
Consistency 9.2 9.2
Tests Correct
Attempt pass rate 60.0% 65.0%
Flaky tests 2 2
Total Runs 60 60
Cost per result 0.397 0.217
Total Cost $0.044 $0.026
Input Price $0.100 / 1M $0.250 / 1M
Output Price $0.400 / 1M $1.500 / 1M
Output Tokens 2,555 2,726
Reasoning Tokens 95,974 8,951
Response Time (avg) 80.22s 1.92s
Response Time (max) 262.83s 5.66s
Response Time (total) 1363.72s 38.45s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 6.6 10.0 50.0% 0 74.75s 360 9,520
Gemini 3.1 Flash Lite 7.3 6.2 75.0% 2 1.84s 1,013 1,548
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 7.1 9.8 50.0% 0 220.48s 464 34,964
Gemini 3.1 Flash Lite 6.8 10.0 50.0% 0 1.71s 465 763
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 10.0 10.0 100.0% 0 262.83s 404 29,806
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 4.48s 348 975
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 10.0 10.0 100.0% 0 24.27s 246 2,743
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.44s 291 697
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 3.0 10.0 0.0% 0 0ms 0 0
Gemini 3.1 Flash Lite 5.3 10.0 33.3% 0 1.52s 15 1,214
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 5.1 3.4 33.3% 1 36.65s 213 4,210
Gemini 3.1 Flash Lite 4.0 10.0 0.0% 0 1.37s 69 438
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 10.0 10.0 100.0% 0 17.47s 69 2,050
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.52s 72 760
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 8.2 7.2 88.9% 1 31.79s 527 5,667
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.40s 210 1,191
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 10.0 10.0 100.0% 0 88.68s 222 5,235
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 5.66s 234 945
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Seed-2.0-Mini 3.0 10.0 0.0% 0 56.76s 50 1,779
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 1.46s 9 420

Quick Compare

Switch Comparison Pair