Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

DeepSeek: DeepSeek V4 Flash vs Google: Gemini 3.5 Flash

Last updated at: 2026-05-19

Metric DeepSeek V4 Flash DeepSeek V4 Flash high Release: 2026-04-24 Free Available Gemini 3.5 Flash Gemini 3.5 Flash medium Release: 2026-05-19
Score 7.6 9.2
Rank #53 #5
Reliability 10.0 10.0
Consistency 7.9 10.0
Tests Correct
Attempt pass rate 75.4% 89.5%
Flaky tests 5 0
Total Runs 57 57
Cost per result 0.299 2.307
Total Cost $0.033 $0.393
Input Price $0.112 / 1M $1.500 / 1M
Output Price $0.224 / 1M $9.000 / 1M
Output Tokens 10,281 1,971
Reasoning Tokens 98,830 36,659
Response Time (avg) 45.88s 3.90s
Response Time (max) 218.13s 12.05s
Response Time (total) 871.76s 74.13s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 8.3 10.0 75.0% 0 28.51s 140 7,770
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.09s 171 3,385
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 62.48s 369 9,361
Gemini 3.5 Flash 10.0 10.0 100.0% 0 8.22s 431 5,190
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 76.57s 465 7,347
Gemini 3.5 Flash 10.0 10.0 100.0% 0 12.05s 351 7,807
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 28.03s 201 1,179
Gemini 3.5 Flash 10.0 10.0 100.0% 0 4.07s 279 3,784
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.1 4.4 44.5% 2 100.31s 27 59,249
Gemini 3.5 Flash 7.7 10.0 66.7% 0 5.24s 12 8,047
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.1 3.1 66.7% 1 25.15s 79 632
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.52s 115 1,144
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 15.36s 63 1,622
Gemini 3.5 Flash 9.9 10.0 100.0% 0 2.70s 71 2,855
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.4 4.4 77.8% 2 25.53s 193 2,597
Gemini 3.5 Flash 7.7 10.0 66.7% 0 2.38s 295 2,747
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 74.73s 228 542
Gemini 3.5 Flash 10.0 10.0 100.0% 0 3.81s 234 455
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.0 10.0 0.0% 0 54.46s 8,516 8,531
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.75s 12 1,245

Quick Compare

Switch Comparison Pair