Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 2.5 Flash vs Google: Gemini 3.1 Flash Lite Preview

Last updated at: 2026-03-15

Metric Gemini 2.5 Flash Gemini 2.5 Flash medium Release: 2025-06-17 Gemini 3.1 Flash Lite Preview Gemini 3.1 Flash Lite Preview medium Release: 2026-03-03
Rank #15 #16
Score 8.0 8.0
Consistency 9.5 10.0
Cost per result 2.619 0.443
Total Cost $0.288 $0.049
Tests Correct
Attempt pass rate 72.9% 68.8%
Flaky tests 1 0
Total Runs 48 48
Output Tokens 1,370 1,731
Reasoning Tokens 110,522 25,821
Response Time (avg) 12.35s 3.83s
Response Time (max) 95.48s 14.93s
Response Time (total) 197.62s 61.25s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 7.8 10.0 66.7% 0 6.98s 249 8,832
Gemini 3.1 Flash Lite Preview 8.8 10.0 66.7% 0 2.53s 564 3,780
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 28.44s 303 11,922
Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 14.93s 327 7,347
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 4.06s 279 2,325
Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 2.29s 279 2,952
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 5.9 7.2 55.6% 1 37.34s 18 80,702
Gemini 3.1 Flash Lite Preview 3.0 10.0 0.0% 0 4.21s 18 5,325
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 4.8 10.0 0.0% 0 4.86s 92 1,899
Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 3.16s 96 1,488
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 9.8 10.0 100.0% 0 2.62s 69 1,203
Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 1.91s 72 2,121
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 7.7 10.0 66.7% 0 3.94s 126 2,499
Gemini 3.1 Flash Lite Preview 7.7 10.0 66.7% 0 3.58s 141 1,896
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 6.20s 234 1,140
Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 3.80s 234 912

Quick Compare

Switch Comparison Pair