Qwen3.8 27B (medium) leads on average score with 7.8 vs 7.7. Local hardware cost was not measured, so cost comparisons are unavailable. Qwen3.8 27B (medium) is faster at 33.05s vs 72.24s, with pass rates of 77.3% vs 66.7%.
Last updated at:
2026-08-15
Rank
#73
Total Output Tokens
241,309
Response Time (avg)
72.24s
Total Cost
$0.767
Rank
#68
Total Output Tokens
136,162
Response Time (avg)
33.05s
Total Cost
N/A
Recommended modelQwen3.8 27B (medium)
It has the best score here (7.8), while responding about 2.2x faster than Seed-2.0-Code (low).
Qwen3.8 27BQwen3.8 27BmediumCandidate responses were generated on local hardware. OpenRouter was used only for judging.Release: 2026-08-14
Score
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#73
#68
Reliability
8.5First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
9.6First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
7.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
9.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
66/66
66/66
Tests Correct
A test is fully passed only if every run passed for that test.Wrong answer: 6API error: 2Did not follow instructions: 1Response Time (avg)72.24sResponse Time (max)485.92sResponse Time (total)1589.31sA test is fully passed only if every run passed for that test.…
8.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
91.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)39.56sResponse Time (max)120.09sResponse Time (total)158.26sA test is fully passed only if every run passed for that test.…
39.56sResponse Time (avg)…
930Total Input Tokens…
2,650Output Tokens…
15,987Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.93sResponse Time (max)5.18sResponse Time (total)11.70sA test is fully passed only if every run passed for that test.…
7.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
55.6%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Wrong answer: 1Response Time (avg)109.70sResponse Time (max)141.71sResponse Time (total)329.11sA test is fully passed only if every run passed for that test.…
109.70sResponse Time (avg)…
7,948Total Input Tokens…
455Output Tokens…
55,759Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)40.50sResponse Time (max)72.14sResponse Time (total)121.51sA test is fully passed only if every run passed for that test.…
7.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)99.12sResponse Time (max)173.03sResponse Time (total)198.23sA test is fully passed only if every run passed for that test.…
99.12sResponse Time (avg)…
55,969Total Input Tokens…
3,774Output Tokens…
21,408Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
8.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Response Time (avg)124.48sResponse Time (max)214.63sResponse Time (total)248.96sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)24.88sResponse Time (max)30.28sResponse Time (total)49.75sA test is fully passed only if every run passed for that test.…
24.88sResponse Time (avg)…
8,046Total Input Tokens…
246Output Tokens…
1,639Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)8.76sResponse Time (max)10.36sResponse Time (total)17.53sA test is fully passed only if every run passed for that test.…
4.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.5%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)238.45sResponse Time (max)485.92sResponse Time (total)715.34sA test is fully passed only if every run passed for that test.…
238.45sResponse Time (avg)…
975Total Input Tokens…
16Output Tokens…
132,626Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)76.98sResponse Time (max)144.07sResponse Time (total)230.93sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)11.14sResponse Time (max)11.14sResponse Time (total)11.14sA test is fully passed only if every run passed for that test.…
11.14sResponse Time (avg)…
579Total Input Tokens…
179Output Tokens…
1,117Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)9.20sResponse Time (max)9.20sResponse Time (total)9.20sA test is fully passed only if every run passed for that test.…
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)12.20sResponse Time (max)21.22sResponse Time (total)24.41sA test is fully passed only if every run passed for that test.…
12.20sResponse Time (avg)…
828Total Input Tokens…
72Output Tokens…
1,034Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.54sResponse Time (max)2.67sResponse Time (total)5.07sA test is fully passed only if every run passed for that test.…
9.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)12.50sResponse Time (max)20.68sResponse Time (total)37.50sA test is fully passed only if every run passed for that test.…
12.50sResponse Time (avg)…
885Total Input Tokens…
445Output Tokens…
2,470Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
8.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)12.62sResponse Time (max)29.62sResponse Time (total)37.85sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)55.56sResponse Time (max)55.56sResponse Time (total)55.56sA test is fully passed only if every run passed for that test.…
55.56sResponse Time (avg)…
9,315Total Input Tokens…
296Output Tokens…
637Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)10.02sResponse Time (max)10.02sResponse Time (total)10.02sA test is fully passed only if every run passed for that test.…
10.02sResponse Time (avg)…
182Total Input Tokens…
9Output Tokens…
490Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)11.29sResponse Time (max)11.29sResponse Time (total)11.29sA test is fully passed only if every run passed for that test.…