The average score is effectively tied at 8.6 vs 8.7. Muse Spark 1.2 (high) has the lower benchmark cost at $1.771 vs $3.004. Claude Fable 5 (medium) is faster at 15.01s vs 26.59s, with pass rates of 78.8% vs 74.2%.
Last updated at:
2026-09-18
Compared models
Rank
#48
Total Output Tokens
42,146
Response Time (avg)
15.01s
Total Cost
$3.004
Rank
#45
Total Output Tokens
382,087
Response Time (avg)
26.59s
Total Cost
$1.771
Recommended modelMuse Spark 1.2 (high)
It has the best score here (8.7), while costing about 1.7x less than Claude Fable 5 (medium).
8.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#48
#45
Reliability
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
9.9First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
9.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
8.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
66/66
66/66
Tests Correct
A test is fully passed only if every run passed for that test.No answer: 3Invalid tool call: 1Wrong answer: 1Response Time (avg)15.01sResponse Time (max)80.80sResponse Time (total)330.15sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 4Did not follow instructions: 3Invalid tool call: 1Response Time (avg)26.59sResponse Time (max)192.00sResponse Time (total)584.89sA test is fully passed only if every run passed for that test.…
Attempt pass rate
78.8%Attempt pass rate = passed attempts / total attempts across runs.…
74.2%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
5Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
17.669
12.645
Total Cost
$3.004
$1.771
Input Price
$10.000 / 1MInput Price…
$1.250 / 1MInput Price…
Output Price
$50.000 / 1MOutput Price…
$4.250 / 1MOutput Price…
Total Input Tokens
89,640Total Input Tokens…
117,037Total Input Tokens…
Output Tokens
33,092Output Tokens…
6,983Output Tokens…
Reasoning Tokens
9,054Reasoning Tokens…
375,104Reasoning Tokens…
Response Time (avg)
15.01sResponse Time (avg)…
26.59sResponse Time (avg)…
Response Time (max)
80.80sResponse Time (max)…
192.00sResponse Time (max)…
Response Time (total)
330.15sResponse Time (total)…
584.89sResponse Time (total)…
Parameters
~9.5T total (~878B active)
~1T total (~50B active)
Availability
Closed
Closed
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)6.20sResponse Time (max)6.57sResponse Time (total)24.81sA test is fully passed only if every run passed for that test.…
7.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
58.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 2Response Time (avg)6.00sResponse Time (max)10.62sResponse Time (total)23.99sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)15.59sResponse Time (max)22.20sResponse Time (total)46.76sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)22.24sResponse Time (max)28.05sResponse Time (total)66.72sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Response Time (avg)27.47sResponse Time (max)33.70sResponse Time (total)54.94sA test is fully passed only if every run passed for that test.…
8.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Response Time (avg)39.89sResponse Time (max)52.46sResponse Time (total)79.78sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)7.18sResponse Time (max)7.44sResponse Time (total)14.35sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)6.93sResponse Time (max)8.10sResponse Time (total)13.85sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Wrong answer: 1Response Time (avg)37.32sResponse Time (max)80.80sResponse Time (total)111.95sA test is fully passed only if every run passed for that test.…
4.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.5%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)108.85sResponse Time (max)192.00sResponse Time (total)326.56sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)7.42sResponse Time (max)7.42sResponse Time (total)7.42sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)13.93sResponse Time (max)13.93sResponse Time (total)13.93sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)5.90sResponse Time (max)5.96sResponse Time (total)11.79sA test is fully passed only if every run passed for that test.…
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)5.43sResponse Time (max)6.96sResponse Time (total)10.86sA test is fully passed only if every run passed for that test.…
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)5.18sResponse Time (max)5.58sResponse Time (total)15.53sA test is fully passed only if every run passed for that test.…
8.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
77.8%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)10.07sResponse Time (max)18.77sResponse Time (total)30.22sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)16.96sResponse Time (max)16.96sResponse Time (total)16.96sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)4.00sResponse Time (max)4.00sResponse Time (total)4.00sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)25.64sResponse Time (max)25.64sResponse Time (total)25.64sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)14.99sResponse Time (max)14.99sResponse Time (total)14.99sA test is fully passed only if every run passed for that test.…