AI BENCHY Compare

DeepSeek: DeepSeek V4 Pro vs Google: Gemma 4 31B

Summary

DeepSeek V4 Pro vs Gemma 4 31B benchmark comparison: The average score is effectively tied at 6.3 vs 6.3. Gemma 4 31B has the lower benchmark cost at $0.033 vs $0.079. Gemma 4 31B is faster at 56.55s vs 65.21s, with pass rates of 52.4% vs 69.8%.

Recommended model: Gemma 4 31B - It has the best score here (6.3), while costing about 2.4x less than DeepSeek V4 Pro.

Last updated at: 2026-06-12

Metric	DeepSeek V4 Pro DeepSeek V4 Pro high Release: 2026-04-24	Gemma 4 31B Gemma 4 31B medium Release: 2026-04-02 Free Available

Metric	DeepSeek V4 Pro DeepSeek V4 Pro high Release: 2026-04-24	Gemma 4 31B Gemma 4 31B medium Release: 2026-04-02 Free Available
Score	6.3	6.3
Rank	#90	#87
Reliability	9.0	10.0
Consistency	7.6	9.4
Tests Correct
Attempt pass rate	52.4%	69.8%
Flaky tests	6	1
Total Runs	63	63
Cost per result	2.869	0.257
Total Cost	$0.079	$0.033
Input Price	$0.435 / 1M	$0.120 / 1M
Output Price	$0.870 / 1M	$0.350 / 1M
Total Input Tokens	32,240	17,957
Output Tokens	12,250	22,356
Reasoning Tokens	72,257	65,726
Response Time (avg)	65.21s	56.55s
Response Time (max)	358.35s	437.40s
Response Time (total)	1304.19s	1074.41s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#90 DeepSeek V4 Pro

high

Cost: $0.023
Time: 257.6s
Tokens: 14,870 tok

#87 Gemma 4 31B

medium

Cost: $0.002
Time: 45.7s
Tokens: 2,696 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	6.4	7.9	58.3%	1		16.53s	448	71	3,617
Gemma 4 31B	10.0	10.0	100.0%	0		12.89s	816	962	2,046

Coding	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	3.3	6.4	11.1%	1		118.23s	1,966	111	20,940
Gemma 4 31B	4.3	5.8	22.2%	1		219.76s	5,568	11,098	33,212

Combined	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	10.0	10.0	100.0%	0		65.02s	14,016	465	5,914
Gemma 4 31B	3.0	10.0	0.0%	0		0ms	0	0	0

Data parsing and extraction	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	7.3	5.9	83.3%	1		23.62s	5,633	229	1,710
Gemma 4 31B	10.0	10.0	100.0%	0		21.11s	8,334	1,822	2,951

Domain specific	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	2.9	7.2	11.1%	1		205.66s	430	10,529	28,089
Gemma 4 31B	7.7	10.0	66.7%	0		38.48s	876	4,349	8,985

General Intelligence	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	6.1	3.1	66.7%	1		25.09s	314	76	1,152
Gemma 4 31B	10.0	10.0	100.0%	0		9.57s	567	105	888

Instructions following	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	10.0	10.0	100.0%	0		41.16s	627	205	2,416
Gemma 4 31B	10.0	10.0	100.0%	0		12.76s	777	533	2,035

Puzzle Solving	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	5.9	7.2	55.6%	1		34.84s	544	139	4,019
Gemma 4 31B	9.9	10.0	100.0%	0		26.91s	801	1,795	5,595

Tool Calling	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	10.0	10.0	100.0%	0		21.33s	8,079	372	593
Gemma 4 31B	3.0	10.0	0.0%	0		0ms	0	0	0

Trivia	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	3.0	10.0	0.0%	0		39.14s	183	53	3,807
Gemma 4 31B	3.0	10.0	0.0%	0		90.14s	218	1,692	10,014

Quick Compare

Switch Comparison Pair

DeepSeek V4 ProhighvsGPT-5.5none DeepSeek V4 ProhighvsQwen3.5-35B-A3Bmedium Gemma 4 31BmediumFree AvailablevsGPT-5.5none DeepSeek V4 ProhighvsNemotron 3 SupermediumFree Available DeepSeek V4 PrononevsGemma 4 31BmediumFree Available Seed-2.0-LitenonevsDeepSeek V4 Prohigh Seed-2.0-LitenonevsGemma 4 31BmediumFree Available DeepSeek V4 ProhighvsGemini 2.5 Flashnone DeepSeek V4 ProhighvsGemini 3.1 Flash Liteminimal DeepSeek V4 ProhighvsGemini 3.1 Flash Litenone DeepSeek V4 ProhighvsGemini 3.1 Flash Litelow DeepSeek V4 ProhighvsQwen3.5-Flashnone