AI BENCHY Compare

DeepSeek: DeepSeek V4 Pro vs Google: Gemma 4 26B A4B

Summary

DeepSeek V4 Pro vs Gemma 4 26B A4B benchmark comparison: The average score is effectively tied at 7.2 vs 7.2. DeepSeek V4 Pro has the lower benchmark cost at $0.030 vs $0.045. DeepSeek V4 Pro is faster at 5.30s vs 63.41s, with pass rates of 52.4% vs 69.8%.

Recommended model: DeepSeek V4 Pro - It has the best score here (7.2), while costing about 1.5x less than Gemma 4 26B A4B.

Last updated at: 2026-06-12

Metric	DeepSeek V4 Pro DeepSeek V4 Pro none Release: 2026-04-24	Gemma 4 26B A4B Gemma 4 26B A4B medium Release: 2026-04-03 Free Available

Metric	DeepSeek V4 Pro DeepSeek V4 Pro none Release: 2026-04-24	Gemma 4 26B A4B Gemma 4 26B A4B medium Release: 2026-04-03 Free Available
Score	7.2	7.2
Rank	#61	#62
Reliability	9.9	10.0
Consistency	8.8	9.2
Tests Correct
Attempt pass rate	52.4%	69.8%
Flaky tests	3	2
Total Runs	61	63
Cost per result	0.293	0.361
Total Cost	$0.030	$0.045
Input Price	$0.435 / 1M	$0.060 / 1M
Output Price	$0.870 / 1M	$0.330 / 1M
Total Input Tokens	53,078	40,252
Output Tokens	7,047	28,000
Reasoning Tokens	0	100,490
Response Time (avg)	5.30s	63.41s
Response Time (max)	23.74s	369.32s
Response Time (total)	111.39s	1268.28s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#61 DeepSeek V4 Pro

none

Invalid SVG

Cost: $0.000
Time: 300.0s
Tokens: 0 tok

#62 Gemma 4 26B A4B

medium

Invalid SVG

Cost: $0.000
Time: 300.0s
Tokens: 0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	3.2	6.1	16.7%	2		4.02s	540	1,168	0
Gemma 4 26B A4B	10.0	10.0	100.0%	0		6.20s	816	1,142	3,045

Coding	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	5.6	10.0	33.3%	0		5.62s	6,795	1,123	0
Gemma 4 26B A4B	2.9	10.0	0.0%	0		272.54s	5,062	14,838	44,567

Combined	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	9.5	10.0	100.0%	0		23.74s	27,529	2,235	0
Gemma 4 26B A4B	9.6	10.0	100.0%	0		73.55s	17,092	5,415	13,112

Data parsing and extraction	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	10.0	10.0	100.0%	0		4.61s	7,568	200	0
Gemma 4 26B A4B	10.0	10.0	100.0%	0		16.51s	8,334	1,567	2,827

Domain specific	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	5.3	10.0	33.3%	0		3.72s	666	24	0
Gemma 4 26B A4B	2.9	4.4	22.2%	2		23.62s	516	2,469	7,105

General Intelligence	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	5.0	10.0	0.0%	0		2.05s	471	126	0
Gemma 4 26B A4B	10.0	10.0	100.0%	0		29.76s	567	25	5,075

Instructions following	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	6.3	5.8	66.7%	1		4.12s	627	713	0
Gemma 4 26B A4B	10.0	10.0	100.0%	0		17.54s	777	887	4,470

Puzzle Solving	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	10.0	10.0	100.0%	0		3.61s	594	442	0
Gemma 4 26B A4B	10.0	10.0	100.0%	0		5.79s	801	410	2,128

Tool Calling	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	10.0	10.0	100.0%	0		7.40s	8,105	328	0
Gemma 4 26B A4B	10.0	10.0	100.0%	0		9.01s	6,096	450	1,256

Trivia	Score	Consistency	Attempt pass rate	Flaky tests	Tests Correct	Response Time (avg)	Input Tokens	Output Tokens	Reasoning Tokens
DeepSeek V4 Pro	3.0	10.0	0.0%	0		5.76s	183	688	0
Gemma 4 26B A4B	3.0	10.0	0.0%	0		180.87s	191	797	16,905

Quick Compare

Switch Comparison Pair

Gemma 4 26B A4BmediumFree AvailablevsQwen3.7 Plusnone Claude Opus 4.8nonevsGemma 4 26B A4BmediumFree Available Gemma 4 26B A4BmediumFree AvailablevsStep 3.7 Flashhigh DeepSeek V4 PrononevsGLM 5V Turbomedium DeepSeek V4 PrononevsMiMo-V2-Flashmedium DeepSeek V4 PrononevsStep 3.7 Flashhigh DeepSeek V4 PrononevsGLM 5.1medium Claude Sonnet 4.6nonevsGemma 4 26B A4BmediumFree Available DeepSeek V4 PrononevsKimi K2.7 Codemedium DeepSeek V4 PrononevsGrok 4.20medium DeepSeek V4 PrononevsGemini 3 Flash Previewlow DeepSeek V4 PrononevsStep 3.5 Flashmedium