Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 3.5 Flash vs OpenAI: GPT-5.5

Last updated at: 2026-05-19

Metric Gemini 3.5 Flash Gemini 3.5 Flash none Release: 2026-05-19 GPT-5.5 GPT-5.5 low Release: 2026-04-24
Score 9.1 8.9
Rank #6 #10
Reliability 10.0 10.0
Consistency 9.0 10.0
Tests Correct
Attempt pass rate 91.7% 84.2%
Flaky tests 2 0
Total Runs 57 57
Cost per result 3.490 4.412
Total Cost $0.489 $0.706
Input Price $1.500 / 1M $5.000 / 1M
Output Price $9.000 / 1M $30.000 / 1M
Output Tokens 53,202 2,008
Reasoning Tokens 0 16,914
Response Time (avg) 5.59s 8.80s
Response Time (max) 14.88s 56.19s
Response Time (total) 89.50s 167.26s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.53s 5,101 0
GPT-5.5 10.0 10.0 100.0% 0 4.43s 246 1,011
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 14.88s 11,611 0
GPT-5.5 10.0 10.0 100.0% 0 7.79s 369 936
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 0.0 0.0 0.0% 0 0ms 0 0
GPT-5.5 10.0 10.0 100.0% 0 9.56s 303 717
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 8.10s 5,895 0
GPT-5.5 10.0 10.0 100.0% 0 3.28s 228 157
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 7.6 7.2 77.8% 1 10.64s 17,910 0
GPT-5.5 5.3 10.0 33.3% 0 27.57s 69 11,731
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 3.46s 1,620 0
GPT-5.5 10.0 10.0 100.0% 0 7.14s 146 170
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 9.8 10.0 100.0% 0 3.38s 3,928 0
GPT-5.5 9.9 10.0 100.0% 0 2.98s 93 356
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 3.13s 4,640 0
GPT-5.5 10.0 10.0 100.0% 0 4.94s 274 895
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 0.0 0.0 0.0% 0 0ms 0 0
GPT-5.5 10.0 10.0 100.0% 0 4.96s 250 101
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.5 Flash 2.8 1.6 33.3% 1 4.87s 2,497 0
GPT-5.5 3.0 10.0 0.0% 0 10.06s 30 840

Quick Compare

Switch Comparison Pair