Qwen3.8 27B leads on average score with 5.5 vs 5.4. Local hardware cost was not measured, so cost comparisons are unavailable. Qwen3.8 27B is faster at 3.12s vs 111.60s, with pass rates of 27.3% vs 40.9%.
Last updated at:
2026-08-15
Rank
#220
Total Output Tokens
584,864
Response Time (avg)
111.60s
Total Cost
$0.114
Rank
#211
Total Output Tokens
11,445
Response Time (avg)
3.12s
Total Cost
N/A
Recommended modelQwen3.8 27B
It has the best score here (5.5), while responding about 35.7x faster than Laguna S 2.1 (high).
A test is fully passed only if every run passed for that test.Wrong answer: 11Did not follow instructions: 1Invalid tool call: 1Response Time (avg)3.12sResponse Time (max)44.76sResponse Time (total)68.69sA test is fully passed only if every run passed for that test.…
Attempt pass rate
27.3%Attempt pass rate = passed attempts / total attempts across runs.…
40.9%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
5Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
3.153Shows the average cost per correct benchmark answer in cents (lower is better).…
N/AShows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$0.114Total Cost (Current Price)…
N/ATotal Cost (Current Price)…
Input Price
$0.090 / 1MInput Price…
N/AInput Price…
Output Price
$0.180 / 1MOutput Price…
N/AOutput Price…
Total Input Tokens
208,937Total Input Tokens…
122,463Total Input Tokens…
Output Tokens
175,588Output Tokens…
11,445Output Tokens…
Reasoning Tokens
409,276Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
111.60sResponse Time (avg)…
3.12sResponse Time (avg)…
Response Time (max)
1190.11sResponse Time (max)…
44.76sResponse Time (max)…
Response Time (total)
2455.24sResponse Time (total)…
68.69sResponse Time (total)…
Parameters
118B total (8B active)
27.3B
Availability
Weights available
Weights available
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#220 Laguna S 2.1
high
Cost
$0.001
Time
6.8s
Tokens
1,389 tok
#211 Qwen3.8 27B
noneCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3No answer: 1Response Time (avg)52.20sResponse Time (max)207.32sResponse Time (total)208.79sA test is fully passed only if every run passed for that test.…
52.20sResponse Time (avg)…
672Total Input Tokens…
504Output Tokens…
46,579Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)828msResponse Time (max)1.95sResponse Time (total)3.31sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1No answer: 1Response Time (avg)271.42sResponse Time (max)423.72sResponse Time (total)814.25sA test is fully passed only if every run passed for that test.…
271.42sResponse Time (avg)…
6,496Total Input Tokens…
59,355Output Tokens…
151,009Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)1.19sResponse Time (max)1.97sResponse Time (total)3.56sA test is fully passed only if every run passed for that test.…
2.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
16.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 2Response Time (avg)702.26sResponse Time (max)1190.11sResponse Time (total)1404.52sA test is fully passed only if every run passed for that test.…
702.26sResponse Time (avg)…
184,420Total Input Tokens…
114,577Output Tokens…
203,826Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Wrong answer: 1Response Time (avg)24.10sResponse Time (max)44.76sResponse Time (total)48.21sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.96sResponse Time (max)3.41sResponse Time (total)3.91sA test is fully passed only if every run passed for that test.…
1.96sResponse Time (avg)…
7,689Total Input Tokens…
243Output Tokens…
1,705Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.14sResponse Time (max)1.18sResponse Time (total)2.27sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)2.51sResponse Time (max)6.82sResponse Time (total)7.52sA test is fully passed only if every run passed for that test.…
2.51sResponse Time (avg)…
771Total Input Tokens…
24Output Tokens…
2,035Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)552msResponse Time (max)675msResponse Time (total)1.66sA test is fully passed only if every run passed for that test.…
4.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)855msResponse Time (max)855msResponse Time (total)855msA test is fully passed only if every run passed for that test.…
855msResponse Time (avg)…
513Total Input Tokens…
135Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
4.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.31sResponse Time (max)1.31sResponse Time (total)1.31sA test is fully passed only if every run passed for that test.…
4.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)453msResponse Time (max)492msResponse Time (total)905msA test is fully passed only if every run passed for that test.…
453msResponse Time (avg)…
702Total Input Tokens…
60Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)582msResponse Time (max)633msResponse Time (total)1.16sA test is fully passed only if every run passed for that test.…
2.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
11.1%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Extra formatting: 1Response Time (avg)1.62sResponse Time (max)3.81sResponse Time (total)4.87sA test is fully passed only if every run passed for that test.…
1.62sResponse Time (avg)…
699Total Input Tokens…
370Output Tokens…
1,130Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)1.25sResponse Time (max)1.84sResponse Time (total)3.76sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.69sResponse Time (max)2.69sResponse Time (total)2.69sA test is fully passed only if every run passed for that test.…
2.69sResponse Time (avg)…
6,747Total Input Tokens…
306Output Tokens…
334Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.75sResponse Time (max)2.75sResponse Time (total)2.75sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)6.94sResponse Time (max)6.94sResponse Time (total)6.94sA test is fully passed only if every run passed for that test.…
6.94sResponse Time (avg)…
228Total Input Tokens…
14Output Tokens…
2,658Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)703msResponse Time (max)703msResponse Time (total)703msA test is fully passed only if every run passed for that test.…