Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

DeepSeek: DeepSeek V4 Flash vs Google: Gemini 2.5 Flash

Last updated at: 2026-04-24

Metric DeepSeek V4 Flash DeepSeek V4 Flash high Release: 2026-04-24 Gemini 2.5 Flash Gemini 2.5 Flash medium Release: 2025-06-17
Score 7.8 8.2
Rank #37 #17
Reliability N/A N/A
Consistency 7.8 9.5
Tests Correct
Attempt pass rate 79.6% 75.9%
Flaky tests 5 1
Total Runs 52 54
Cost per result 0.189 2.454
Total Cost $0.021 $0.319
Input Price $0.140 / 1M $0.300 / 1M
Output Price $0.280 / 1M $2.500 / 1M
Output Tokens 1,757 1,898
Reasoning Tokens 55,907 122,273
Response Time (avg) 47.47s 12.12s
Response Time (max) 255.28s 95.48s
Response Time (total) 854.45s 218.12s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 8.3 10.0 75.0% 0 28.51s 140 7,770
Gemini 2.5 Flash 8.4 10.0 75.0% 0 6.30s 255 10,233
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 62.48s 369 9,361
Gemini 2.5 Flash 10.0 10.0 100.0% 0 16.23s 522 10,350
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 76.57s 465 7,347
Gemini 2.5 Flash 10.0 10.0 100.0% 0 28.44s 303 11,922
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 28.03s 201 1,179
Gemini 2.5 Flash 10.0 10.0 100.0% 0 4.06s 279 2,325
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.1 4.4 44.5% 2 112.69s 19 24,857
Gemini 2.5 Flash 5.9 7.2 55.6% 1 37.34s 18 80,702
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.1 3.1 66.7% 1 25.15s 79 632
Gemini 2.5 Flash 4.8 10.0 0.0% 0 4.86s 92 1,899
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 15.36s 63 1,622
Gemini 2.5 Flash 9.8 10.0 100.0% 0 2.62s 69 1,203
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.4 4.5 77.8% 2 25.53s 193 2,597
Gemini 2.5 Flash 7.7 10.0 66.7% 0 3.94s 126 2,499
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 74.73s 228 542
Gemini 2.5 Flash 10.0 10.0 100.0% 0 6.20s 234 1,140

Quick Compare

Switch Comparison Pair