Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 3.5 Flash vs Inception: Mercury 2

Last updated at: 2026-06-03

Metric Gemini 3.5 Flash Gemini 3.5 Flash low Release: 2026-05-19 Mercury 2 Mercury 2 medium Release: 2026-02-24
Score 9.3 6.5
Rank #3 #89
Reliability 10.0 10.0
Consistency 10.0 8.8
Tests Correct
Attempt pass rate 90.0% 51.7%
Flaky tests 0 3
Total Runs 60 60
Cost per result 1.582 0.611
Total Cost $0.285 $0.055
Input Price $1.500 / 1M $0.250 / 1M
Output Price $9.000 / 1M $0.750 / 1M
Total Input Tokens 33,935 32,570
Output Tokens 2,027 4,022
Reasoning Tokens 23,938 58,405
Response Time (avg) 2.98s 2.27s
Response Time (max) 6.44s 14.63s
Response Time (total) 59.59s 43.20s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.52s 494 209 2,536
Mercury 2 6.9 9.9 50.0% 0 1.12s 554 2,546 2,609
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 6.8 10.0 50.0% 0 5.54s 5,115 452 6,839
Mercury 2 7.2 6.5 66.7% 1 2.29s 4,519 270 8,514
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 6.44s 12,873 351 3,050
Mercury 2 10.0 10.0 100.0% 0 3.28s 12,909 268 4,887
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 1.81s 7,548 279 1,164
Mercury 2 7.3 5.9 83.3% 1 1.11s 6,234 183 1,656
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 7.7 10.0 66.7% 0 3.39s 633 12 4,538
Mercury 2 2.9 7.2 11.1% 1 6.48s 695 41 30,754
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.27s 486 119 916
Mercury 2 4.8 10.0 0.0% 0 821ms 456 137 542
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 9.9 10.0 100.0% 0 1.86s 615 71 1,652
Mercury 2 10.0 10.0 100.0% 0 1.07s 340 14 958
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.35s 558 288 2,150
Mercury 2 5.4 10.0 33.3% 0 949ms 601 361 2,781
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 3.27s 5,457 234 403
Mercury 2 10.0 10.0 100.0% 0 1.89s 6,080 180 1,956
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash 10.0 10.0 100.0% 0 1.88s 156 12 690
Mercury 2 3.0 10.0 0.0% 0 2.58s 182 22 3,748

Quick Compare

Switch Comparison Pair