The average score is effectively tied at 7.9 vs 7.9. Local hardware cost was not measured, so cost comparisons are unavailable. Qwen3.8 27B (low) is faster at 39.11s vs 142.84s, with pass rates of 77.3% vs 68.2%.
Last updated at:
2026-08-15
Rank
#58
Total Output Tokens
224,128
Response Time (avg)
142.84s
Total Cost
$3.467
Rank
#61
Total Output Tokens
169,329
Response Time (avg)
39.11s
Total Cost
N/A
Recommended modelQwen3.8 27B (low)
It has the best score here (7.9), while responding about 3.7x faster than Kimi K3 (max).
Qwen3.8 27BQwen3.8 27BlowCandidate responses were generated on local hardware. OpenRouter was used only for judging.Release: 2026-08-14
Score
7.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#58
#61
Reliability
9.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
9.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
66/66
66/66
Tests Correct
A test is fully passed only if every run passed for that test.API error: 2No answer: 2Extra formatting: 1Timed out: 1Response Time (avg)142.84sResponse Time (max)766.58sResponse Time (total)2714.02sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 4No answer: 2Did not follow instructions: 1Response Time (avg)39.11sResponse Time (max)376.04sResponse Time (total)860.51sA test is fully passed only if every run passed for that test.…
Attempt pass rate
77.3%Attempt pass rate = passed attempts / total attempts across runs.…
68.2%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
21.668Shows the average cost per correct benchmark answer in cents (lower is better).…
N/AShows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$3.467Total Cost (Current Price)…
N/ATotal Cost (Current Price)…
Input Price
$3.000 / 1MInput Price…
N/AInput Price…
Output Price
$15.000 / 1MOutput Price…
N/AOutput Price…
Total Input Tokens
34,960Total Input Tokens…
99,705Total Input Tokens…
Output Tokens
2,887Output Tokens…
1,545Output Tokens…
Reasoning Tokens
221,241Reasoning Tokens…
167,784Reasoning Tokens…
Response Time (avg)
142.84sResponse Time (avg)…
39.11sResponse Time (avg)…
Response Time (max)
766.58sResponse Time (max)…
376.04sResponse Time (max)…
Response Time (total)
2714.02sResponse Time (total)…
860.51sResponse Time (total)…
Parameters
2.8T total (104B active)
27.3B
Availability
Weights available
Weights available
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#58 MoonshotAI: Kimi K3
max
Cost
$0.281
Time
506.9s
Tokens
18,863 tok
#61 Qwen3.8 27B
lowCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)10.24sResponse Time (max)15.21sResponse Time (total)40.97sA test is fully passed only if every run passed for that test.…
10.24sResponse Time (avg)…
1,638Total Input Tokens…
496Output Tokens…
2,629Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.02sResponse Time (max)5.68sResponse Time (total)12.07sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)325.79sResponse Time (max)589.43sResponse Time (total)977.36sA test is fully passed only if every run passed for that test.…
325.79sResponse Time (avg)…
8,061Total Input Tokens…
460Output Tokens…
84,012Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)140.92sResponse Time (max)376.04sResponse Time (total)422.75sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)223.03sResponse Time (max)223.03sResponse Time (total)223.03sA test is fully passed only if every run passed for that test.…
223.03sResponse Time (avg)…
12,792Total Input Tokens…
399Output Tokens…
20,460Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)90.63sResponse Time (max)143.71sResponse Time (total)181.27sA test is fully passed only if every run passed for that test.…
7.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Extra formatting: 1Response Time (avg)16.72sResponse Time (max)18.73sResponse Time (total)33.44sA test is fully passed only if every run passed for that test.…
16.72sResponse Time (avg)…
7,839Total Input Tokens…
354Output Tokens…
2,606Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)8.03sResponse Time (max)9.09sResponse Time (total)16.06sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Timed out: 1Response Time (avg)683.62sResponse Time (max)766.58sResponse Time (total)1367.25sA test is fully passed only if every run passed for that test.…
683.62sResponse Time (avg)…
838Total Input Tokens…
57Output Tokens…
106,892Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)60.78sResponse Time (max)120.94sResponse Time (total)182.34sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)14.91sResponse Time (max)14.91sResponse Time (total)14.91sA test is fully passed only if every run passed for that test.…
14.91sResponse Time (avg)…
732Total Input Tokens…
191Output Tokens…
871Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)5.37sResponse Time (max)5.37sResponse Time (total)5.37sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)7.66sResponse Time (max)7.99sResponse Time (total)15.33sA test is fully passed only if every run passed for that test.…
7.66sResponse Time (avg)…
1,179Total Input Tokens…
147Output Tokens…
1,119Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.43sResponse Time (max)2.87sResponse Time (total)4.86sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)7.36sResponse Time (max)12.30sResponse Time (total)22.07sA test is fully passed only if every run passed for that test.…
7.36sResponse Time (avg)…
1,416Total Input Tokens…
495Output Tokens…
1,368Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)4.91sResponse Time (max)8.81sResponse Time (total)14.74sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0msResponse Time (avg)…
0Total Input Tokens…
0Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)8.42sResponse Time (max)8.42sResponse Time (total)8.42sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)19.67sResponse Time (max)19.67sResponse Time (total)19.67sA test is fully passed only if every run passed for that test.…
19.67sResponse Time (avg)…
465Total Input Tokens…
288Output Tokens…
1,284Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)12.64sResponse Time (max)12.64sResponse Time (total)12.64sA test is fully passed only if every run passed for that test.…