Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.8 vs Gemini 3 PRO Preview

Last updated at: 2026-05-28

Metric Claude Opus 4.8 Claude Opus 4.8 medium Release: 2026-05-28 Gemini 3 PRO Preview Gemini 3 PRO Preview medium Release: 2025-11-18
Score 8.7 8.1
Rank #12 #21
Reliability 10.0 N/A
Consistency 9.6 10.0
Tests Correct
Attempt pass rate 83.3% 73.7%
Flaky tests 1 0
Total Runs 60 60
Cost per result 6.285 1.406
Total Cost $1.006 $0.385
Input Price $5.000 / 1M $9.506 / 1M
Output Price $25.000 / 1M $9.506 / 1M
Output Tokens 23,201 1,490
Reasoning Tokens 5,901 10,102
Response Time (avg) 9.34s 9.05s
Response Time (max) 38.03s 26.24s
Response Time (total) 186.84s 90.53s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 1,179 478
Gemini 3 PRO Preview 10.0 10.0 100.0% 0 14.99s 149 1,485
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 14.97s 6,651 1,381
Gemini 3 PRO Preview 3.0 10.0 0.0% 0 0ms 0 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 9.8 10.0 100.0% 0 38.03s 5,260 1,588
Gemini 3 PRO Preview 3.0 10.0 0.0% 0 10.37s 351 952
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 7.1 5.6 83.3% 1 12.29s 481 312
Gemini 3 PRO Preview 10.0 10.0 100.0% 0 10.84s 279 3,156
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 5.3 10.0 33.3% 0 14.15s 7,477 900
Gemini 3 PRO Preview 5.3 10.0 33.3% 0 7.01s 15 1,195
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 2.46s 237 0
Gemini 3 PRO Preview 10.0 10.0 100.0% 0 9.34s 78 374
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.32s 373 320
Gemini 3 PRO Preview 9.8 10.0 100.0% 0 3.26s 69 754
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 791 483
Gemini 3 PRO Preview 10.0 10.0 100.0% 0 3.88s 225 1,215
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 10.0 10.0 100.0% 0 8.96s 301 225
Gemini 3 PRO Preview 10.0 10.0 100.0% 0 11.96s 324 971
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.8 3.0 10.0 0.0% 0 6.14s 451 214
Gemini 3 PRO Preview 0.0 0.0 0.0% 0 0ms 0 0

Quick Compare

Switch Comparison Pair