Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemma 4 31B vs OpenAI: GPT-5.4 Nano

Summary

Gemma 4 31B vs GPT-5.4 Nano benchmark comparison: GPT-5.4 Nano leads on average score with 7.5 vs 6.1. Gemma 4 31B has the lower benchmark cost at $0.004 vs $0.107. Gemma 4 31B is faster at 4.05s vs 11.95s, with pass rates of 47.6% vs 63.5%.

Recommended model: Gemma 4 31B - It offers the best overall trade-off: a competitive score (6.1), lower cost than GPT-5.4 Nano, and balanced response time.

Last updated at: 2026-06-18

Metric Gemma 4 31B Gemma 4 31B none Release: 2026-04-02 Free Available GPT-5.4 Nano GPT-5.4 Nano medium Release: 2026-03-17
Score 6.1 7.5
Rank #98 #46
Reliability 10.0 10.0
Consistency 10.0 8.4
Tests Correct
Attempt pass rate 47.6% 63.5%
Flaky tests 0 4
Total Runs 63 63
Cost per result 0.034 0.969
Total Cost $0.004 $0.107
Input Price $0.120 / 1M $0.200 / 1M
Output Price $0.350 / 1M $1.250 / 1M
Total Input Tokens 20,911 35,434
Output Tokens 1,407 3,014
Reasoning Tokens 0 76,520
Response Time (avg) 4.05s 11.95s
Response Time (max) 26.13s 94.06s
Response Time (total) 76.87s 250.98s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#98 Gemma 4 31B

none
Cost
$0.001
Time
12.8s
Tokens
795 tok

#46 GPT-5.4 Nano

medium
Cost
$0.007
Time
24.6s
Tokens
4,943 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 1.85s 852 45 0
GPT-5.4 Nano 8.3 10.0 75.0% 0 4.52s 606 683 2,254
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 5.5 10.0 33.3% 0 11.19s 8,381 735 0
GPT-5.4 Nano 6.1 4.7 66.7% 2 19.12s 7,305 516 20,778
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0 0
GPT-5.4 Nano 9.8 10.0 100.0% 0 24.13s 12,345 349 5,719
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.25s 8,352 285 0
GPT-5.4 Nano 10.0 10.0 100.0% 0 2.54s 7,140 234 516
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 7.7 10.0 66.7% 0 3.22s 903 27 0
GPT-5.4 Nano 5.9 7.2 55.6% 1 38.18s 619 60 43,325
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.09s 576 117 0
GPT-5.4 Nano 4.5 10.0 0.0% 0 4.15s 477 179 443
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 2.84s 795 78 0
GPT-5.4 Nano 9.8 10.0 100.0% 0 1.88s 660 95 521
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 33.3% 0 4.23s 828 108 0
GPT-5.4 Nano 4.1 7.2 22.2% 1 3.79s 642 594 1,408
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0 0
GPT-5.4 Nano 10.0 10.0 100.0% 0 7.71s 5,445 234 382
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 1.25s 224 12 0
GPT-5.4 Nano 3.0 10.0 0.0% 0 4.81s 195 70 1,174

Quick Compare

Switch Comparison Pair