Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Claude Sonnet 5.5 (max) vs Muse Spark 1.3 (high)

Claude Sonnet 5.5 (max) leads on average score with 9.0 vs 8.9. Muse Spark 1.3 (high) has the lower benchmark cost at $1.653 vs $7.545. Muse Spark 1.3 (high) is faster at 52.49s vs 83.04s, with pass rates of 81.8% vs 77.3%.

Last updated at: 2026-09-29

Compared models

Rank
#39
Total Output Tokens
727,349
Response Time (avg)
83.04s
Total Cost
$7.545
Rank
#44
Total Output Tokens
359,608
Response Time (avg)
52.49s
Total Cost
$1.653
Recommended model Muse Spark 1.3 (high)

Its score stays close to the best score here (8.9 vs 9.0), while costing about 4.6x less than Claude Sonnet 5.5 (max).

Detailed comparison

Metric Claude Sonnet 5.5 Claude Sonnet 5.5 max Release: 2026-09-29 Muse Spark 1.3 Muse Spark 1.3 high Release: 2026-09-02
Score 9.0 8.9
Rank #39 #44
Reliability 10.0 10.0
Consistency 9.4 9.3
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 81.8% 77.3%
Flaky tests 2 2
Total Runs 66 66
Cost per result 44.380 10.327
Total Cost $7.545 $1.653
Input Price $2.000 / 1M $1.250 / 1M
Output Price $10.000 / 1M $4.250 / 1M
Total Input Tokens 135,496 99,143
Output Tokens 8,153 12,083
Reasoning Tokens 719,196 347,525
Response Time (avg) 83.04s 52.49s
Response Time (max) 488.39s 251.53s
Response Time (total) 1826.84s 1154.76s
Parameters ~1T total (~100B active) ~1T total (~50B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#39 Claude Sonnet 5.5

max
OpenRouter returned no showcase content (finish_reason: length).
Cost
$0.000
Time
571.7s
Tokens
0 tok

#44 Muse Spark 1.3

high
Cost
$0.050
Time
203.3s
Tokens
11,679 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5.5 10.0 10.0 100.0% 0 55.26s 10,608 580 70,730
Muse Spark 1.3 10.0 10.0 100.0% 0 51.66s 7,275 1,228 64,287

Quick Compare

Switch Comparison Pair