Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

Google: Gemma 4 31B vs OpenAI: GPT-5.2

Last updated at: 2026-04-02

Metric Gemma 4 31B Gemma 4 31B none Release: 2026-04-02 GPT-5.2 GPT-5.2 medium Release: 2025-12-11
Score 6.7 7.3
Rank #47 #36
Consistency 10.0 8.0
Tests Correct
Attempt pass rate 52.9% 70.6%
Flaky tests 0 4
Total Runs 51 51
Cost per result 0.023 3.131
Total Cost $0.002 $0.314
Input Price $0.140 / 1M $1.750 / 1M
Output Price $0.400 / 1M $14.000 / 1M
Output Tokens 660 2,238
Reasoning Tokens 0 16,811
Response Time (avg) 2.55s 13.93s
Response Time (max) 4.68s 77.80s
Response Time (total) 38.20s 139.29s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 1.85s 45 0
GPT-5.2 6.5 8.0 58.3% 1 7.81s 567 2,002
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
GPT-5.2 10.0 10.0 100.0% 0 14.06s 291 1,757
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.25s 285 0
GPT-5.2 10.0 10.0 100.0% 0 3.15s 234 420
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 7.7 10.0 66.7% 0 3.22s 27 0
GPT-5.2 5.9 7.2 55.6% 1 77.80s 42 10,342
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.09s 117 0
GPT-5.2 3.7 9.7 0.0% 0 4.32s 162 269
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 2.84s 78 0
GPT-5.2 9.9 10.0 100.0% 0 3.12s 94 614
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 5.5 10.0 33.3% 0 2.95s 108 0
GPT-5.2 7.7 7.3 77.8% 1 5.47s 609 938
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
GPT-5.2 4.7 1.6 66.7% 1 10.30s 239 469

Quick Compare

Switch Comparison Pair