Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemini 3.1 Flash Lite vs OpenAI: GPT-5.4 Mini

Last updated at: 2026-05-22

Metric Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite low Release: 2026-05-08 GPT-5.4 Mini GPT-5.4 Mini medium Release: 2026-03-17
Score 7.4 7.1
Rank #50 #65
Reliability 10.0 10.0
Consistency 9.2 7.6
Tests Correct
Attempt pass rate 65.0% 68.3%
Flaky tests 2 6
Total Runs 60 60
Cost per result 0.217 4.867
Total Cost $0.026 $0.487
Input Price $0.250 / 1M $0.750 / 1M
Output Price $1.500 / 1M $4.500 / 1M
Output Tokens 2,726 2,186
Reasoning Tokens 8,951 100,706
Response Time (avg) 1.92s 22.14s
Response Time (max) 5.66s 138.75s
Response Time (total) 38.45s 442.74s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 7.3 6.2 75.0% 2 1.84s 1,013 1,548
GPT-5.4 Mini 8.6 7.9 91.7% 1 4.05s 296 2,876
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 6.8 10.0 50.0% 0 1.71s 465 763
GPT-5.4 Mini 7.5 6.0 83.3% 1 73.25s 446 32,513
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 4.48s 348 975
GPT-5.4 Mini 10.0 10.0 100.0% 0 17.81s 317 4,317
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.44s 291 697
GPT-5.4 Mini 10.0 10.0 100.0% 0 2.43s 234 650
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 5.3 10.0 33.3% 0 1.52s 15 1,214
GPT-5.4 Mini 4.1 4.4 44.5% 2 65.31s 60 43,286
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 4.0 10.0 0.0% 0 1.37s 69 438
GPT-5.4 Mini 4.5 10.0 0.0% 0 3.72s 150 510
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.52s 72 760
GPT-5.4 Mini 7.4 6.7 66.7% 1 2.50s 129 1,337
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.40s 210 1,191
GPT-5.4 Mini 7.8 10.0 66.7% 0 4.33s 271 2,449
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 5.66s 234 945
GPT-5.4 Mini 4.7 1.6 66.7% 1 9.62s 251 2,594
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 1.46s 9 420
GPT-5.4 Mini 3.0 10.0 0.0% 0 30.10s 32 10,174

Quick Compare

Switch Comparison Pair