Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

Google: Gemini 3 Flash Preview vs Grok 4.20 Beta

Last updated at: 2026-04-26

Metric Gemini 3 Flash Preview Gemini 3 Flash Preview medium Release: 2025-12-17 Grok 4.20 Beta Grok 4.20 Beta none Release: 2026-03-12
Score 10.0 5.3
Rank #1 #93
Reliability N/A N/A
Consistency 10.0 9.2
Tests Correct
Attempt pass rate 100.0% 29.6%
Flaky tests 0 2
Total Runs 18 52
Cost per result 0.600 2.255
Total Cost $0.108 $0.091
Input Price $0.500 / 1M $0.000 / 1M
Output Price $3.000 / 1M $0.000 / 1M
Output Tokens 655 1,591
Reasoning Tokens 33,749 0
Response Time (avg) 12.11s 1.19s
Response Time (max) 82.37s 6.48s
Response Time (total) 217.93s 21.37s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 3.26s 110 1,076
Grok 4.20 Beta 4.0 8.4 16.7% 1 597ms 251 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 82.37s 144 16,257
Grok 4.20 Beta 5.5 10.0 0.0% 0 1.14s 74 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 23.58s 117 3,495
Grok 4.20 Beta 3.0 10.0 0.0% 0 6.48s 282 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 7.62s 93 2,197
Grok 4.20 Beta 10.0 10.0 100.0% 0 601ms 197 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 14.81s 4 7,228
Grok 4.20 Beta 3.0 10.0 0.0% 0 611ms 160 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 6.34s 24 635
Grok 4.20 Beta 5.0 10.0 0.0% 0 541ms 87 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 4.30s 24 903
Grok 4.20 Beta 4.8 10.0 0.0% 0 687ms 60 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 4.86s 61 1,455
Grok 4.20 Beta 5.9 7.2 55.6% 1 541ms 291 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 3 Flash Preview 10.0 10.0 100.0% 0 9.78s 78 503
Grok 4.20 Beta 10.0 10.0 100.0% 0 4.79s 189 0

Quick Compare

Switch Comparison Pair