Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

DeepSeek: DeepSeek V3.2 vs Google: Gemini 2.5 Flash

Last updated at: 2026-04-29

Metric DeepSeek V3.2 DeepSeek V3.2 none Release: 2025-12-01 Gemini 2.5 Flash Gemini 2.5 Flash medium Release: 2025-06-17
Score 6.0 8.2
Rank #84 #20
Reliability N/A N/A
Consistency 8.6 9.5
Tests Correct
Attempt pass rate 46.3% 75.9%
Flaky tests 3 1
Total Runs 52 54
Cost per result 0.225 2.454
Total Cost $0.016 $0.319
Input Price $0.252 / 1M $0.300 / 1M
Output Price $0.378 / 1M $2.500 / 1M
Output Tokens 8,378 1,898
Reasoning Tokens 0 122,273
Response Time (avg) 12.07s 12.12s
Response Time (max) 115.89s 95.48s
Response Time (total) 217.28s 218.12s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 3.2 9.8 0.0% 0 7.63s 1,419 0
Gemini 2.5 Flash 8.4 10.0 75.0% 0 6.30s 255 10,233
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 2.4 1.3 33.3% 1 7.63s 553 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 16.23s 522 10,350
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 6.5 10.0 0.0% 0 115.89s 2,887 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 28.44s 303 11,922
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 6.3 5.8 66.7% 1 9.42s 1,710 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 4.06s 279 2,325
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 3.0 10.0 0.0% 0 1.52s 18 0
Gemini 2.5 Flash 5.9 7.2 55.6% 1 37.34s 18 80,702
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 2.86s 67 0
Gemini 2.5 Flash 4.8 10.0 0.0% 0 4.86s 92 1,899
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 1.52s 66 0
Gemini 2.5 Flash 9.8 10.0 100.0% 0 2.62s 69 1,203
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 8.5 7.5 88.9% 1 7.37s 1,136 0
Gemini 2.5 Flash 7.7 10.0 66.7% 0 3.94s 126 2,499
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 11.85s 522 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 6.20s 234 1,140

Quick Compare

Switch Comparison Pair