Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

DeepSeek: DeepSeek V4 Flash vs Google: Gemini 3.1 Flash Lite

Last updated at: 2026-05-08

Metric DeepSeek V4 Flash DeepSeek V4 Flash high Release: 2026-04-24 Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite medium Release: 2026-05-08
Score 7.6 7.9
Rank #48 #27
Reliability 10.0 10.0
Consistency 7.9 9.1
Tests Correct
Attempt pass rate 75.4% 71.9%
Flaky tests 5 2
Total Runs 57 57
Cost per result 0.299 0.452
Total Cost $0.033 $0.059
Input Price $0.140 / 1M $0.250 / 1M
Output Price $0.280 / 1M $1.500 / 1M
Output Tokens 10,281 2,224
Reasoning Tokens 98,830 32,034
Response Time (avg) 45.88s 3.14s
Response Time (max) 218.13s 10.87s
Response Time (total) 871.76s 59.62s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 8.3 10.0 75.0% 0 28.51s 140 7,770
Gemini 3.1 Flash Lite 9.1 10.0 75.0% 0 2.39s 604 4,201
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 62.48s 369 9,361
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 3.26s 429 2,712
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 76.57s 465 7,347
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 10.87s 327 7,401
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 28.03s 201 1,179
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 2.60s 279 2,845
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.1 4.4 44.5% 2 100.31s 27 59,249
Gemini 3.1 Flash Lite 2.9 7.2 11.1% 1 3.16s 15 5,165
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.1 3.1 66.7% 1 25.15s 79 632
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 2.60s 84 1,142
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 15.36s 63 1,622
Gemini 3.1 Flash Lite 9.9 10.0 100.0% 0 2.59s 75 3,320
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.4 4.4 77.8% 2 25.53s 193 2,597
Gemini 3.1 Flash Lite 7.6 7.2 77.8% 1 1.95s 165 2,450
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 74.73s 228 542
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 4.55s 234 921
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.0 10.0 0.0% 0 54.46s 8,516 8,531
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 3.08s 12 1,877

Quick Compare

Switch Comparison Pair