Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.7 vs OpenAI: GPT-5.5

Summary

Claude Opus 4.7 vs GPT-5.5 benchmark comparison: GPT-5.5 leads on average score with 9.0 vs 7.4. Claude Opus 4.7 has the lower benchmark cost at $0.505 vs $3.679. Claude Opus 4.7 is faster at 3.02s vs 37.98s, with pass rates of 76.2% vs 87.3%.

Recommended model: Claude Opus 4.7 - It offers the best overall trade-off: a competitive score (7.4), lower cost than GPT-5.5, and balanced response time.

Last updated at: 2026-06-18

Metric Claude Opus 4.7 Claude Opus 4.7 none Release: 2026-04-16 GPT-5.5 GPT-5.5 medium Release: 2026-04-24
Score 7.4 9.0
Rank #49 #9
Reliability 10.0 10.0
Consistency 9.0 8.9
Tests Correct
Attempt pass rate 76.2% 87.3%
Flaky tests 0 3
Total Runs 57 63
Cost per result 3.154 21.638
Total Cost $0.505 $3.679
Input Price $5.000 / 1M $5.000 / 1M
Output Price $25.000 / 1M $30.000 / 1M
Total Input Tokens 69,576 34,212
Output Tokens 6,265 1,985
Reasoning Tokens 0 114,925
Response Time (avg) 3.02s 37.98s
Response Time (max) 18.27s 332.10s
Response Time (total) 57.44s 797.60s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#49 Claude Opus 4.7

none
Cost
$0.051
Time
24.2s
Tokens
2,181 tok

#9 GPT-5.5

medium
Cost
$0.112
Time
71.9s
Tokens
3,807 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 8.3 10.0 75.0% 0 2.12s 894 522 0
GPT-5.5 10.0 10.0 100.0% 0 4.66s 606 250 1,335
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 3.3 3.3 33.3% 0 2.84s 1,176 494 0
GPT-5.5 8.8 7.8 88.9% 1 59.77s 7,305 362 24,959
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 9.5 10.0 100.0% 0 18.27s 37,740 3,504 0
GPT-5.5 10.0 10.0 100.0% 0 19.29s 11,019 312 2,841
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.15s 10,533 324 0
GPT-5.5 10.0 10.0 100.0% 0 4.18s 7,140 234 593
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 7.7 10.0 66.7% 0 1.19s 1,020 78 0
GPT-5.5 5.3 7.2 44.4% 1 164.14s 723 67 79,625
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 3.47s 723 257 0
GPT-5.5 10.0 10.0 100.0% 0 4.16s 477 138 223
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 1.46s 939 114 0
GPT-5.5 10.0 10.0 100.0% 0 3.36s 660 93 538
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 2.46s 939 597 0
GPT-5.5 10.0 10.0 100.0% 0 6.76s 642 241 2,225
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 10.0 10.0 100.0% 0 4.74s 15,339 372 0
GPT-5.5 10.0 10.0 100.0% 0 10.57s 5,445 258 832
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.7 3.0 10.0 0.0% 0 1.46s 273 3 0
GPT-5.5 2.8 1.6 33.3% 1 37.86s 195 30 1,754

Quick Compare

Switch Comparison Pair