Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemini 2.5 Flash vs Google: Gemma 4 31B

Last updated at: 2026-06-01

Metric Gemini 2.5 Flash Gemini 2.5 Flash none Release: 2025-06-17 Gemma 4 31B Gemma 4 31B none Release: 2026-04-02 Free Available
Score 6.4 6.7
Rank #95 #83
Reliability 10.0 10.0
Consistency 9.6 10.0
Tests Correct
Attempt pass rate 48.3% 50.0%
Flaky tests 1 0
Total Runs 60 60
Cost per result 0.159 0.030
Total Cost $0.015 $0.003
Input Price $0.300 / 1M $0.120 / 1M
Output Price $2.500 / 1M $0.370 / 1M
Output Tokens 1,764 1,398
Reasoning Tokens 0 0
Response Time (avg) 889ms 4.05s
Response Time (max) 4.39s 26.13s
Response Time (total) 17.79s 72.97s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 3.0 10.0 0.0% 0 582ms 102 0
Gemma 4 31B 6.5 10.0 50.0% 0 1.85s 45 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 6.8 10.0 50.0% 0 810ms 477 0
Gemma 4 31B 6.8 10.0 50.0% 0 14.84s 726 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 3.0 10.0 0.0% 0 4.39s 366 0
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 652ms 279 0
Gemma 4 31B 10.0 10.0 100.0% 0 2.25s 285 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 5.9 7.2 55.6% 1 495ms 12 0
Gemma 4 31B 7.7 10.0 66.7% 0 3.22s 27 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 5.0 10.0 0.0% 0 615ms 78 0
Gemma 4 31B 10.0 10.0 100.0% 0 2.09s 117 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 590ms 72 0
Gemma 4 31B 6.5 10.0 50.0% 0 2.84s 78 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 7.7 10.0 66.7% 0 604ms 132 0
Gemma 4 31B 6.5 10.0 33.3% 0 4.23s 108 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 1.91s 234 0
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 3.0 10.0 0.0% 0 1.15s 12 0
Gemma 4 31B 3.0 10.0 0.0% 0 1.25s 12 0

Quick Compare

Switch Comparison Pair