Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Google: Gemini 3.1 Flash Lite vs xAI: Grok Build 0.1

Last updated at: 2026-05-22

Metric Gemini 3.1 Flash Lite Gemini 3.1 Flash Lite minimal Release: 2026-05-08 Grok Build 0.1 Grok Build 0.1 none Release: 2026-05-21
Score 6.7 6.6
Rank #78 #82
Reliability 10.0 10.0
Consistency 8.8 8.0
Tests Correct
Attempt pass rate 56.7% 60.4%
Flaky tests 3 4
Total Runs 60 57
Cost per result 0.123 7.805
Total Cost $0.013 $0.547
Input Price $0.250 / 1M $1.000 / 1M
Output Price $1.500 / 1M $2.000 / 1M
Output Tokens 2,481 267,275
Reasoning Tokens 0 0
Response Time (avg) 1.37s 28.69s
Response Time (max) 4.49s 138.35s
Response Time (total) 27.32s 459.00s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 8.3 10.0 75.0% 0 1.10s 639 0
Grok Build 0.1 8.7 7.9 91.7% 1 6.30s 11,162 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 6.8 10.0 50.0% 0 951ms 660 0
Grok Build 0.1 10.0 10.0 100.0% 0 21.41s 16,568 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 2.53s 357 0
Grok Build 0.1 0.0 0.0 0.0% 0 0ms 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 1.04s 279 0
Grok Build 0.1 4.7 1.6 66.7% 1 9.33s 6,359 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 2.9 7.2 11.1% 1 1.02s 15 0
Grok Build 0.1 3.6 7.2 22.2% 1 103.71s 179,469 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 4.0 10.0 0.0% 0 791ms 63 0
Grok Build 0.1 4.3 10.0 0.0% 0 12.47s 6,647 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 932ms 72 0
Grok Build 0.1 9.8 10.0 100.0% 0 7.36s 8,970 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 6.0 4.6 66.7% 2 2.15s 153 0
Grok Build 0.1 6.4 7.7 55.6% 1 9.55s 14,982 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 10.0 10.0 100.0% 0 3.51s 234 0
Grok Build 0.1 0.0 0.0 0.0% 0 0ms 0 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3.1 Flash Lite 3.0 10.0 0.0% 0 724ms 9 0
Grok Build 0.1 3.0 10.0 0.0% 0 36.09s 23,118 0

Quick Compare

Switch Comparison Pair