Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 3 Flash Preview vs Google: Gemma 4 31B

Last updated at: 2026-05-19

Metric Gemini 3 Flash Preview Gemini 3 Flash Preview low Release: 2025-12-17 Gemma 4 31B Gemma 4 31B medium Release: 2026-04-02 Free Available
Score 8.8 8.2
Rank #11 #18
Reliability 10.0 6.7
Consistency 9.6 9.6
Tests Correct
Attempt pass rate 86.0% 77.2%
Flaky tests 1 1
Total Runs 57 57
Cost per result 0.579 0.158
Total Cost $0.093 $0.023
Input Price $0.500 / 1M $0.120 / 1M
Output Price $3.000 / 1M $0.370 / 1M
Output Tokens 2,027 14,426
Reasoning Tokens 23,906 37,964
Response Time (avg) 5.84s 28.72s
Response Time (max) 14.72s 90.14s
Response Time (total) 110.87s 488.27s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 3.48s 281 3,082
Gemma 4 31B 10.0 10.0 100.0% 0 12.89s 962 2,046
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 6.94s 426 2,717
Gemma 4 31B 4.7 1.6 66.7% 1 70.97s 3,166 5,449
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 3.0 10.0 0.0% 0 3.27s 326 0
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 9.40s 279 3,656
Gemma 4 31B 10.0 10.0 100.0% 0 21.11s 1,822 2,951
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 5.3 7.2 44.4% 1 8.05s 12 6,410
Gemma 4 31B 7.7 10.0 66.7% 0 38.48s 4,349 8,985
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 3.68s 120 981
Gemma 4 31B 10.0 10.0 100.0% 0 9.57s 105 888
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 9.9 10.0 100.0% 0 7.02s 71 2,752
Gemma 4 31B 10.0 10.0 100.0% 0 12.76s 533 2,035
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 6.11s 269 3,260
Gemma 4 31B 9.9 10.0 100.0% 0 27.63s 1,797 5,596
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 4.99s 234 415
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 2.75s 9 633
Gemma 4 31B 3.0 10.0 0.0% 0 90.14s 1,692 10,014

Quick Compare

Switch Comparison Pair