Navigate
Advertise here

DeepSeek V4 Pro 0423 vs Gemini 3.5 Flash Lite (medium)

DeepSeek V4 Pro 0423 leads on average score with 7.4 vs 7.3. DeepSeek V4 Pro 0423 has the lower benchmark cost at $0.083 vs $0.440. Gemini 3.5 Flash Lite (medium) is faster at 8.68s vs 14.05s, with pass rates of 49.3% vs 76.8%.

Last updated at: 2026-10-01

Compared models

Rank
#129
Total Output Tokens
46,004
Response Time (avg)
14.05s
Total Cost
$0.083
Rank
#138
Total Output Tokens
129,214
Response Time (avg)
8.68s
Total Cost
$0.440
Recommended model DeepSeek V4 Pro 0423

It has the best score here (7.4), while costing about 5.3x less than Gemini 3.5 Flash Lite (medium).

Detailed comparison

Metric DeepSeek V4 Pro 0423 DeepSeek V4 Pro 0423 none Release: 2026-04-24 Gemini 3.5 Flash Lite Gemini 3.5 Flash Lite medium Release: 2026-07-21
Score 7.4 7.3
Rank #129 #138
Reliability 10.0 10.0
Consistency 8.7 7.6
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 49.3% 76.8%
Flaky tests 4 7
Total Runs 69 69
Cost per result 2.341 3.140
Total Cost $0.083 $0.440
Input Price $0.209 / 1M $0.300 / 1M
Output Price $0.418 / 1M $2.500 / 1M
Cache Read Price $0.018 / 1M $0.030 / 1M
Cache Write Price N/A $0.084 / 1M
Total Input Tokens 304,186 388,307
Output Tokens 46,004 10,412
Reasoning Tokens 0 118,802
Response Time (avg) 14.05s 8.68s
Response Time (max) 119.44s 86.92s
Response Time (total) 323.17s 199.65s
Parameters 1.6T total (49B active) ~150B total (~10B active)
Availability Open source Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#129 DeepSeek V4 Pro 0423

none
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

#138 Gemini 3.5 Flash Lite

medium
Cost
$0.010
Time
17.8s
Tokens
4,000 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 0423 5.6 10.0 33.3% 0 13.38s 7,275 5,500 0
Gemini 3.5 Flash Lite 7.9 9.9 66.7% 0 9.96s 8,122 471 28,930

Quick Compare

Switch Comparison Pair