Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

DeepSeek: DeepSeek V3.2 vs DeepSeek: DeepSeek V4 Pro

Last updated at: 2026-05-29

Metric DeepSeek V3.2 DeepSeek V3.2 none Release: 2025-12-01 DeepSeek V4 Pro DeepSeek V4 Pro high Release: 2026-04-24
Score 6.2 7.0
Rank #97 #79
Reliability 10.0 8.9
Consistency 8.3 8.7
Tests Correct
Attempt pass rate 48.3% 63.3%
Flaky tests 4 3
Total Runs 60 60
Cost per result 0.222 1.935
Total Cost $0.018 $0.213
Input Price $0.252 / 1M $0.435 / 1M
Output Price $0.378 / 1M $0.870 / 1M
Output Tokens 11,159 12,244
Reasoning Tokens 0 53,958
Response Time (avg) 14.43s 58.92s
Response Time (max) 115.89s 358.35s
Response Time (total) 288.55s 1119.51s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 3.8 8.2 12.5% 1 9.35s 1,073 0
DeepSeek V4 Pro 8.3 10.0 75.0% 0 16.53s 71 3,617
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 3.1 5.4 16.7% 1 20.87s 4,522 0
DeepSeek V4 Pro 3.0 5.0 25.0% 1 51.77s 105 2,641
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 6.5 10.0 0.0% 0 115.89s 2,887 0
DeepSeek V4 Pro 10.0 10.0 100.0% 0 65.02s 465 5,914
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 6.3 5.8 66.7% 1 9.42s 1,710 0
DeepSeek V4 Pro 10.0 10.0 100.0% 0 23.62s 229 1,710
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 3.2 6.9 16.7% 1 4.17s 21 0
DeepSeek V4 Pro 3.2 6.9 16.7% 1 205.66s 10,529 28,089
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 9.32s 43 0
DeepSeek V4 Pro 6.1 3.1 66.7% 1 25.09s 76 1,152
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 1.52s 66 0
DeepSeek V4 Pro 10.0 10.0 100.0% 0 41.16s 205 2,416
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 6.91s 298 0
DeepSeek V4 Pro 7.7 10.0 66.7% 0 34.84s 139 4,019
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 10.0 10.0 100.0% 0 11.85s 522 0
DeepSeek V4 Pro 10.0 10.0 100.0% 0 21.33s 372 593
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
DeepSeek V3.2 3.0 10.0 0.0% 0 17.23s 17 0
DeepSeek V4 Pro 3.0 10.0 0.0% 0 39.14s 53 3,807

Quick Compare

Switch Comparison Pair