Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemini 3.5 Flash vs OpenAI: GPT-5.4

Last updated at: 2026-06-04

Metric Gemini 3.5 Flash Gemini 3.5 Flash medium Release: 2026-05-19 GPT-5.4 GPT-5.4 medium Release: 2026-03-05
Score 9.0 8.0
Rank #7 #21
Reliability 10.0 10.0
Consistency 9.6 8.6
Tests Correct
Attempt pass rate 87.3% 76.2%
Flaky tests 1 4
Total Runs 63 63
Cost per result 3.229 8.640
Total Cost $0.582 $1.210
Input Price $1.500 / 1M $2.500 / 1M
Output Price $9.000 / 1M $15.000 / 1M
Total Input Tokens 36,936 34,108
Output Tokens 2,001 2,242
Reasoning Tokens 56,408 72,707
Response Time (avg) 4.94s 22.35s
Response Time (max) 18.07s 100.41s
Response Time (total) 103.79s 469.29s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.09s 492 171 3,385
GPT-5.4 8.3 10.0 75.0% 0 4.11s 606 240 1,511
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 7.9 7.5 77.8% 1 12.63s 8,118 461 24,939
GPT-5.4 8.8 7.8 88.9% 1 44.36s 7,305 433 24,216
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 12.05s 12,873 351 7,807
GPT-5.4 10.0 10.0 100.0% 0 20.57s 11,019 301 3,543
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 4.07s 7,548 279 3,784
GPT-5.4 10.0 10.0 100.0% 0 5.32s 7,140 234 804
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 7.7 10.0 66.7% 0 5.24s 633 12 8,047
GPT-5.4 5.3 7.2 44.4% 1 74.27s 619 61 34,748
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.52s 486 115 1,144
GPT-5.4 4.7 3.1 33.3% 1 4.92s 477 145 321
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 9.9 10.0 100.0% 0 2.70s 615 71 2,855
GPT-5.4 10.0 10.0 100.0% 0 3.11s 660 93 897
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 7.7 10.0 66.7% 0 2.38s 558 295 2,747
GPT-5.4 8.2 7.2 88.9% 1 9.14s 642 441 3,815
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 3.81s 5,457 234 455
GPT-5.4 10.0 10.0 100.0% 0 13.28s 5,445 264 1,031
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.75s 156 12 1,245
GPT-5.4 3.0 10.0 0.0% 0 13.95s 195 30 1,821

Quick Compare

Switch Comparison Pair