Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.7 vs Anthropic: Claude Opus 4.8

Last updated at: 2026-05-28

Metric Claude Opus 4.7 Claude Opus 4.7 medium Release: 2026-04-16 Claude Opus 4.8 Claude Opus 4.8 medium Release: 2026-05-28
Score 8.9 8.7
Rank #7 #12
Reliability 10.0 10.0
Consistency 10.0 9.6
Tests Correct
Attempt pass rate 85.0% 83.3%
Flaky tests 0 1
Total Runs 60 60
Cost per result 3.670 6.285
Total Cost $0.624 $1.006
Input Price $5.000 / 1M $5.000 / 1M
Output Price $25.000 / 1M $25.000 / 1M
Output Tokens 10,439 23,201
Reasoning Tokens 2,198 5,901
Response Time (avg) 4.48s 9.34s
Response Time (max) 23.18s 38.03s
Response Time (total) 85.21s 186.84s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 8.3 10.0 75.0% 0 1.85s 348 0
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 1,179 478
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 14.79s 6,210 1,114
Claude Opus 4.8 10.0 10.0 100.0% 0 14.97s 6,651 1,381
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 21.45s 2,369 1,084
Claude Opus 4.8 9.8 10.0 100.0% 0 38.03s 5,260 1,588
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.37s 324 0
Claude Opus 4.8 7.1 5.6 83.3% 1 12.29s 481 312
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 7.7 10.0 66.7% 0 1.17s 51 0
Claude Opus 4.8 5.3 10.0 33.3% 0 14.15s 7,477 900
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.87s 256 0
Claude Opus 4.8 10.0 10.0 100.0% 0 2.46s 237 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 1.57s 114 0
Claude Opus 4.8 10.0 10.0 100.0% 0 3.32s 373 320
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.43s 370 0
Claude Opus 4.8 10.0 10.0 100.0% 0 3.95s 791 483
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 4.17s 373 0
Claude Opus 4.8 10.0 10.0 100.0% 0 8.96s 301 225
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.7 3.0 10.0 0.0% 0 2.25s 24 0
Claude Opus 4.8 3.0 10.0 0.0% 0 6.14s 451 214

Quick Compare

Switch Comparison Pair