Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Sonnet 5 vs Google: Gemini 2.5 Flash

Summary

Claude Sonnet 5 vs Gemini 2.5 Flash benchmark comparison: Gemini 2.5 Flash leads on average score with 6.2 vs 5.7. Gemini 2.5 Flash has the lower benchmark cost at $0.016 vs $0.287. Gemini 2.5 Flash is faster at 875ms vs 4.74s, with pass rates of 42.9% vs 46.0%.

Recommended model: Gemini 2.5 Flash - It has the best score here (6.2), while costing about 18.9x less than Claude Sonnet 5.

Last updated at: 2026-06-30

Metric Claude Sonnet 5 Claude Sonnet 5 none Release: 2026-06-30 Gemini 2.5 Flash Gemini 2.5 Flash none Release: 2025-06-17
Score 5.7 6.2
Rank #117 #95
Reliability 10.0 10.0
Consistency 8.6 9.6
Tests Correct
Attempt pass rate 42.9% 46.0%
Flaky tests 4 1
Total Runs 63 63
Cost per result 4.098 0.169
Total Cost $0.287 $0.016
Input Price $2.000 / 1M $0.300 / 1M
Output Price $10.000 / 1M $2.500 / 1M
Total Input Tokens 76,797 35,926
Output Tokens 13,325 1,770
Reasoning Tokens 0 0
Response Time (avg) 4.74s 875ms
Response Time (max) 29.46s 4.39s
Response Time (total) 99.46s 18.37s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#117 Claude Sonnet 5

none
Cost
$0.061
Time
53.7s
Tokens
6,172 tok

#95 Gemini 2.5 Flash

none
Invalid SVG
Cost
$0.164
Time
215.5s
Tokens
65,659 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 5.3 10.0 25.0% 0 3.60s 834 1,813 0
Gemini 2.5 Flash 3.0 10.0 0.0% 0 582ms 492 102 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 4.6 7.9 22.2% 1 3.67s 10,590 1,864 0
Gemini 2.5 Flash 5.5 10.0 33.3% 0 736ms 8,122 483 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 3.0 10.0 0.0% 0 29.46s 38,775 6,340 0
Gemini 2.5 Flash 3.0 10.0 0.0% 0 4.39s 12,519 366 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 10.0 10.0 100.0% 0 3.01s 10,503 309 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 652ms 7,257 279 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 5.3 7.2 44.4% 1 3.28s 975 933 0
Gemini 2.5 Flash 5.9 7.2 55.6% 1 495ms 633 12 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 4.7 3.1 33.3% 1 2.81s 708 272 0
Gemini 2.5 Flash 5.0 10.0 0.0% 0 615ms 486 78 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 6.4 10.0 50.0% 0 2.58s 909 103 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 590ms 615 72 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 6.0 7.4 55.6% 1 3.22s 894 778 0
Gemini 2.5 Flash 7.7 10.0 66.7% 0 604ms 558 132 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 10.0 10.0 100.0% 0 6.80s 12,351 522 0
Gemini 2.5 Flash 10.0 10.0 100.0% 0 1.91s 5,088 234 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 3.0 10.0 0.0% 0 4.31s 258 391 0
Gemini 2.5 Flash 3.0 10.0 0.0% 0 1.15s 156 12 0

Quick Compare

Switch Comparison Pair