Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemma 4 31B vs OpenAI: GPT-5.2

Last updated at: 2026-05-10

Metric Gemma 4 31B Gemma 4 31B none Release: 2026-04-02 Free Available GPT-5.2 GPT-5.2 medium Release: 2025-12-11
Score 6.9 7.2
Rank #66 #60
Reliability 10.0 10.0
Consistency 10.0 8.2
Tests Correct
Attempt pass rate 52.6% 68.4%
Flaky tests 0 4
Total Runs 57 57
Cost per result 0.025 3.609
Total Cost $0.003 $0.397
Input Price $0.130 / 1M $1.750 / 1M
Output Price $0.380 / 1M $14.000 / 1M
Output Tokens 1,371 2,731
Reasoning Tokens 0 22,200
Response Time (avg) 3.86s 15.22s
Response Time (max) 26.13s 77.80s
Response Time (total) 65.57s 182.59s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 1.85s 45 0
GPT-5.2 6.5 8.0 58.3% 1 7.81s 567 2,002
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 26.13s 699 0
GPT-5.2 10.0 10.0 100.0% 0 15.12s 467 2,166
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
GPT-5.2 10.0 10.0 100.0% 0 14.06s 291 1,757
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.25s 285 0
GPT-5.2 10.0 10.0 100.0% 0 3.15s 234 420
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 7.7 10.0 66.7% 0 3.22s 27 0
GPT-5.2 5.9 7.2 55.6% 1 77.80s 42 10,342
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 10.0 10.0 100.0% 0 2.09s 117 0
GPT-5.2 3.7 9.7 0.0% 0 4.32s 162 269
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 50.0% 0 2.84s 78 0
GPT-5.2 9.9 10.0 100.0% 0 3.12s 94 614
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 6.5 10.0 33.3% 0 2.95s 108 0
GPT-5.2 7.6 7.3 77.8% 1 5.47s 609 938
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 0ms 0 0
GPT-5.2 4.7 1.6 66.7% 1 10.30s 239 469
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemma 4 31B 3.0 10.0 0.0% 0 1.25s 12 0
GPT-5.2 3.0 10.0 0.0% 0 28.18s 26 3,223

Quick Compare

Switch Comparison Pair