Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Muse Spark 1.3 (medium) vs Qwen3.8 Max (0902) (high)

Qwen3.8 Max (0902) (high) leads on average score with 8.9 vs 8.8. Muse Spark 1.3 (medium) has the lower benchmark cost at $1.469 vs $2.736. Muse Spark 1.3 (medium) is faster at 38.57s vs 185.56s, with pass rates of 84.9% vs 86.4%.

Last updated at: 2026-09-07

Rank
#37
Total Output Tokens
319,264
Response Time (avg)
38.57s
Total Cost
$1.469
Rank
#31
Total Output Tokens
419,443
Response Time (avg)
185.56s
Total Cost
$2.736
Recommended model Muse Spark 1.3 (medium)

Its score stays close to the best score here (8.8 vs 8.9), while costing about 1.9x less than Qwen3.8 Max (0902) (high).

Detailed comparison

Metric Muse Spark 1.3 Muse Spark 1.3 medium Release: 2026-09-02 Qwen3.8 Max (0902) Qwen3.8 Max (0902) high Release: 2026-09-07
Score 8.8 8.9
Rank #37 #31
Reliability 10.0 9.8
Consistency 8.9 9.0
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 84.9% 86.4%
Flaky tests 3 3
Total Runs 66 66
Cost per result 8.638 16.089
Total Cost $1.469 $2.736
Input Price $1.250 / 1M $2.000 / 1M
Output Price $4.250 / 1M $6.000 / 1M
Total Input Tokens 89,177 109,226
Output Tokens 11,044 6,890
Reasoning Tokens 308,220 412,553
Response Time (avg) 38.57s 185.56s
Response Time (max) 219.77s 912.78s
Response Time (total) 848.59s 4082.41s
Parameters ~1T total (~50B active) 2.4T total (~100B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#37 Muse Spark 1.3

medium
Cost
$0.025
Time
113.4s
Tokens
5,952 tok

#31 Qwen3.8 Max (0902)

high
Reached the allocated time limit (600 seconds) without receiving showcase output.
Cost
$0.000
Time
600.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Muse Spark 1.3 10.0 10.0 100.0% 0 41.14s 7,275 1,182 46,737
Qwen3.8 Max (0902) 10.0 10.0 100.0% 0 253.22s 8,235 418 91,734

Quick Compare

Switch Comparison Pair