Summary
Claude Opus 5.5 scores 1.8 on AI BENCHY and ranks #359. It has 10.0 reliability, a 4.6% pass rate, $0.015 total cost, and 13.58s average response time.
What makes Claude Opus 5.5 unique: Its total benchmark cost is unusually low for its score range.
Archived model: this model is no longer updated or tested on new tests.
Model facts
Researched on 2026-09-23
- Parameters
- ~5T total (~500B active)
- Architecture
- MoE
- Availability
- Closed
- License
- -
Best estimate from public evidence; the vendor did not disclose every value. Vendor does not disclose parameter counts or architecture. Low-confidence catalog estimate carries forward the prior anthropic/claude-opus-5 family estimate; the linked announcement establishes identity and availability, not these counts.
1.8
Consistency
5.0
10.0
$0.015
Total Output Tokens
735
Total Input Tokens
73
Input Price
$4.000 / 1M
Output Price
$20.000 / 1M
Flaky tests
0
Flaky tests had mixed outcomes across runs (at least one pass and one fail).
Charts
Choose the first model, then click a second model to open a side-by-side page.
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Category | Score | Consistency | Tests Correct |
|---|---|---|---|
| Anti-AI Tricks | 0.0 | 0.0 | |
| Coding | 0.0 | 0.0 | |
| Combined | 0.0 | 0.0 | |
| Data parsing and extraction | 0.0 | 0.0 | |
| Domain specific | 3.3 | 3.3 | |
| General Intelligence | 0.0 | 0.0 | |
| Instructions following | 0.0 | 0.0 | |
| Puzzle Solving | 0.0 | 0.0 | |
| Tool Calling | 0.0 | 0.0 | |
| Trivia | 0.0 | 0.0 |