Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

Google: Gemma 4 31B vs OpenAI: GPT-5.3-Codex

Last updated at: 2026-04-02

Metric Gemma 4 31B Gemma 4 31B none Release: 2026-04-02 GPT-5.3-Codex GPT-5.3-Codex medium Release: 2026-02-05
Score 6.7 8.5
Rank #47 #8
Consistency 10.0 8.6
Tests Correct
Attempt pass rate 52.9% 82.4%
Flaky tests 0 3
Total Runs 51 51
Cost per result 0.023 4.526
Total Cost $0.002 $0.544
Input Price $0.140 / 1M $1.750 / 1M
Output Price $0.400 / 1M $14.000 / 1M
Output Tokens 660 1,788
Reasoning Tokens 0 33,649
Response Time (avg) 2.55s 15.76s
Response Time (max) 4.68s 100.93s
Response Time (total) 38.20s 267.97s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 1.85s 45 0
GPT-5.3-Codex 8.7 7.9 91.7% 1 4.16s 240 1,722
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
GPT-5.3-Codex 10.0 10.0 100.0% 0 19.56s 364 2,731
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.25s 285 0
GPT-5.3-Codex 10.0 10.0 100.0% 0 3.07s 234 728
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 7.7 10.0 66.7% 0 3.22s 27 0
GPT-5.3-Codex 5.9 7.2 55.6% 1 64.31s 64 25,308
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.09s 117 0
GPT-5.3-Codex 4.6 10.0 0.0% 0 4.87s 187 331
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 2.84s 78 0
GPT-5.3-Codex 10.0 10.0 100.0% 0 3.04s 93 693
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 5.5 10.0 33.3% 0 2.95s 108 0
GPT-5.3-Codex 9.0 7.9 88.9% 1 5.12s 352 1,644
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
GPT-5.3-Codex 10.0 10.0 100.0% 0 6.37s 254 492

Quick Compare

Switch Comparison Pair