Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Sonnet 4.6 vs Google: Gemini 3.5 Flash

Last updated at: 2026-05-19

Metric Claude Sonnet 4.6 Claude Sonnet 4.6 none Release: 2026-02-17 Gemini 3.5 Flash Gemini 3.5 Flash medium Release: 2026-05-19
Score 7.2 9.2
Rank #68 #5
Reliability 10.0 10.0
Consistency 9.6 10.0
Tests Correct
Attempt pass rate 61.4% 89.5%
Flaky tests 1 0
Total Runs 57 57
Cost per result 2.441 2.307
Total Cost $0.269 $0.393
Input Price $3.000 / 1M $1.500 / 1M
Output Price $15.000 / 1M $9.000 / 1M
Output Tokens 7,864 1,971
Reasoning Tokens 0 36,659
Response Time (avg) 4.96s 3.90s
Response Time (max) 23.84s 12.05s
Response Time (total) 59.50s 74.13s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 4.8 10.0 25.0% 0 2.94s 1,214 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.09s 171 3,385
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 3.67s 523 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 8.22s 431 5,190
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 9.5 10.0 100.0% 0 23.84s 3,766 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 12.05s 351 7,807
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 3.43s 252 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 4.07s 279 3,784
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 7.7 10.0 66.7% 0 3.54s 413 0
Gemini 3.5 Flash 7.7 10.0 66.7% 0 5.24s 12 8,047
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.1 3.1 66.7% 1 2.56s 192 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.52s 115 1,144
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 6.5 10.0 50.0% 0 1.96s 90 0
Gemini 3.5 Flash 9.9 10.0 100.0% 0 2.70s 71 2,855
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 7.7 10.0 66.7% 0 2.92s 536 0
Gemini 3.5 Flash 7.7 10.0 66.7% 0 2.38s 295 2,747
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 10.0 10.0 100.0% 0 4.11s 447 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 3.81s 234 455
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Sonnet 4.6 3.0 10.0 0.0% 0 4.67s 431 0
Gemini 3.5 Flash 10.0 10.0 100.0% 0 2.75s 12 1,245

Quick Compare

Switch Comparison Pair