Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 3.1 Flash Lite vs Inception: Mercury 2

Last updated at: 2026-05-29

Metric Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite minimal Release: 2026-05-08 Mercury 2 Mercury 2 medium Release: 2026-02-24
Score 6.7 6.5
Rank #84 #92
Reliability 10.0 10.0
Consistency 8.8 8.8
Tests Correct
Attempt pass rate 56.7% 51.7%
Flaky tests 3 3
Total Runs 60 60
Cost per result 0.123 0.611
Total Cost $0.013 $0.055
Input Price $0.250 / 1M $0.250 / 1M
Output Price $1.500 / 1M $0.750 / 1M
Output Tokens 2,481 4,022
Reasoning Tokens 0 58,405
Response Time (avg) 1.37s 2.27s
Response Time (max) 4.49s 14.63s
Response Time (total) 27.32s 43.20s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 8.3 10.0 75.0% 0 1.10s 639 0
Mercury 2 6.9 9.9 50.0% 0 1.12s 2,546 2,609
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 6.8 10.0 50.0% 0 951ms 660 0
Mercury 2 7.2 6.5 66.7% 1 2.29s 270 8,514
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 2.53s 357 0
Mercury 2 10.0 10.0 100.0% 0 3.28s 268 4,887
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.04s 279 0
Mercury 2 7.3 5.9 83.3% 1 1.11s 183 1,656
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 2.9 7.2 11.1% 1 1.02s 15 0
Mercury 2 2.9 7.2 11.1% 1 6.48s 41 30,754
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 4.0 10.0 0.0% 0 791ms 63 0
Mercury 2 4.8 10.0 0.0% 0 821ms 137 542
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 932ms 72 0
Mercury 2 10.0 10.0 100.0% 0 1.07s 14 958
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 6.0 4.6 66.7% 2 2.15s 153 0
Mercury 2 5.4 10.0 33.3% 0 949ms 361 2,781
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 3.51s 234 0
Mercury 2 10.0 10.0 100.0% 0 1.89s 180 1,956
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 724ms 9 0
Mercury 2 3.0 10.0 0.0% 0 2.58s 22 3,748

Quick Compare

Switch Comparison Pair