Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemini 3.1 Flash Lite vs OpenAI: GPT-5.2 Chat

Last updated at: 2026-05-08

Metric Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite low Release: 2026-05-08 GPT-5.2 Chat GPT-5.2 Chat none Release: 2025-12-11
Score 7.6 7.6
Rank #44 #41
Reliability 10.0 10.0
Consistency 9.2 8.8
Tests Correct
Attempt pass rate 68.4% 71.9%
Flaky tests 2 3
Total Runs 57 57
Cost per result 0.203 2.572
Total Cost $0.025 $0.309
Input Price $0.250 / 1M $1.750 / 1M
Output Price $1.500 / 1M $14.000 / 1M
Output Tokens 2,702 18,585
Reasoning Tokens 8,596 0
Response Time (avg) 1.92s 6.85s
Response Time (max) 5.66s 38.52s
Response Time (total) 36.49s 130.06s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 7.3 6.2 75.0% 2 1.84s 1,013 1,548
GPT-5.2 Chat 8.7 7.9 91.7% 1 3.40s 1,807 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.46s 441 408
GPT-5.2 Chat 10.0 10.0 100.0% 0 8.97s 1,345 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 4.48s 348 975
GPT-5.2 Chat 10.0 10.0 100.0% 0 9.12s 1,243 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.44s 291 697
GPT-5.2 Chat 10.0 10.0 100.0% 0 3.05s 980 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 5.3 10.0 33.3% 0 1.52s 15 1,214
GPT-5.2 Chat 5.3 10.0 33.3% 0 17.78s 7,810 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 4.0 10.0 0.0% 0 1.37s 69 438
GPT-5.2 Chat 4.4 3.0 33.3% 1 3.20s 335 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.52s 72 760
GPT-5.2 Chat 7.3 5.9 83.3% 1 5.46s 1,528 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.40s 210 1,191
GPT-5.2 Chat 7.7 10.0 66.7% 0 4.42s 1,743 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 5.66s 234 945
GPT-5.2 Chat 10.0 10.0 100.0% 0 4.68s 555 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 1.46s 9 420
GPT-5.2 Chat 3.0 10.0 0.0% 0 6.89s 1,239 0

Quick Compare

Switch Comparison Pair