Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

DeepSeek: DeepSeek V4 Flash vs Google: Gemma 4 31B

Summary

DeepSeek V4 Flash vs Gemma 4 31B benchmark comparison: Gemma 4 31B leads on average score with 6.3 vs 5.5. DeepSeek V4 Flash has the lower benchmark cost at $0.008 vs $0.033. DeepSeek V4 Flash is faster at 26.75s vs 56.55s, with pass rates of 30.2% vs 69.8%.

Recommended model: Gemma 4 31B - It has the strongest score in this comparison (6.3) and the best overall balance of cost and response time across all 2 models.

Last updated at: 2026-06-12

Metric DeepSeek V4 Flash DeepSeek V4 Flash none Release: 2026-04-24 Gemma 4 31B Gemma 4 31B medium Release: 2026-04-02 Free Available
Score 5.5 6.3
Rank #120 #87
Reliability 10.0 10.0
Consistency 8.9 9.4
Tests Correct
Attempt pass rate 30.2% 69.8%
Flaky tests 3 1
Total Runs 63 63
Cost per result 0.203 0.257
Total Cost $0.008 $0.033
Input Price $0.098 / 1M $0.120 / 1M
Output Price $0.196 / 1M $0.350 / 1M
Total Input Tokens 50,127 17,957
Output Tokens 13,710 22,356
Reasoning Tokens 0 65,726
Response Time (avg) 26.75s 56.55s
Response Time (max) 111.96s 437.40s
Response Time (total) 561.82s 1074.41s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#120 DeepSeek V4 Flash

none
Cost
$0.004
Time
157.6s
Tokens
11,297 tok

#87 Gemma 4 31B

medium
Cost
$0.002
Time
45.7s
Tokens
2,696 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.0 10.0 0.0% 0 20.18s 540 174 0
Gemma 4 31B 10.0 10.0 100.0% 0 12.89s 816 962 2,046
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.2 7.4 11.1% 1 17.13s 7,279 9,717 0
Gemma 4 31B 4.3 5.8 22.2% 1 219.76s 5,568 11,098 33,212
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.5 2.1 66.7% 1 111.96s 24,398 2,664 0
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 23.79s 7,290 195 0
Gemma 4 31B 10.0 10.0 100.0% 0 21.11s 8,334 1,822 2,951
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 5.3 10.0 33.3% 0 19.73s 666 18 0
Gemma 4 31B 7.7 10.0 66.7% 0 38.48s 876 4,349 8,985
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 4.2 9.9 0.0% 0 23.74s 471 67 0
Gemma 4 31B 10.0 10.0 100.0% 0 9.57s 567 105 888
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 6.5 10.0 50.0% 0 17.54s 627 321 0
Gemma 4 31B 10.0 10.0 100.0% 0 12.76s 777 533 2,035
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.1 7.3 11.1% 1 23.72s 594 207 0
Gemma 4 31B 9.9 10.0 100.0% 0 26.91s 801 1,795 5,595
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 10.0 10.0 100.0% 0 77.93s 8,079 327 0
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Flash 3.0 10.0 0.0% 0 3.07s 183 20 0
Gemma 4 31B 3.0 10.0 0.0% 0 90.14s 218 1,692 10,014

Quick Compare

Switch Comparison Pair