Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.8 vs Google: Gemini 3.1 Pro Preview

Last updated at: 2026-05-29

Metric Claude Opus 4.8 Claude Opus 4.8 none Release: 2026-05-28 Gemini 3.1 Pro Preview Gemini 3.1 Pro Preview medium Release: 2026-02-19
Score 7.3 9.3
Rank #65 #4
Reliability 10.0 10.0
Consistency 9.2 10.0
Tests Correct
Attempt pass rate 65.0% 90.0%
Flaky tests 2 0
Total Runs 60 60
Cost per result 4.324 5.587
Total Cost $0.519 $1.006
Input Price $5.000 / 1M $2.000 / 1M
Output Price $25.000 / 1M $12.000 / 1M
Output Tokens 8,098 1,971
Reasoning Tokens 0 75,384
Response Time (avg) 3.51s 20.77s
Response Time (max) 17.73s 88.68s
Response Time (total) 70.19s 269.96s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 6.5 10.0 50.0% 0 3.40s 1,472 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 7.90s 112 3,218
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 6.8 10.0 50.0% 0 3.59s 1,323 0
Gemini 3.1 Pro Preview 7.0 9.8 50.0% 0 54.28s 429 37,735
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 9.5 10.0 100.0% 0 17.73s 3,259 0
Gemini 3.1 Pro Preview 9.5 10.0 100.0% 0 40.61s 432 9,281
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 7.3 5.8 83.3% 1 1.77s 308 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 7.72s 279 3,904
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 5.3 7.2 44.4% 1 1.66s 61 0
Gemini 3.1 Pro Preview 7.7 10.0 66.7% 0 32.73s 18 12,424
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.48s 230 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 11.77s 108 1,179
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 9.9 10.0 100.0% 0 1.37s 95 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 9.56s 72 2,236
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 7.7 10.0 66.7% 0 2.74s 783 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 6.90s 235 3,128
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 5.35s 355 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 23.15s 274 982
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 3.0 10.0 0.0% 0 3.41s 212 0
Gemini 3.1 Pro Preview 10.0 10.0 100.0% 0 6.27s 12 1,297

Quick Compare

Switch Comparison Pair