Gemma 4 26B A4B leads on average score with 5.6 vs 5.5. Local hardware cost was not measured, so cost comparisons are unavailable. Qwen3.8 27B is faster at 3.12s vs 7.58s, with pass rates of 43.9% vs 40.9%.
Last updated at:
2026-08-15
Rank
#204
Total Output Tokens
15,781
Response Time (avg)
7.58s
Total Cost
$0.023
Rank
#211
Total Output Tokens
11,445
Response Time (avg)
3.12s
Total Cost
N/A
Recommended modelQwen3.8 27B
Its score stays close to the best score here (5.5 vs 5.6), while responding about 2.4x faster than Gemma 4 26B A4B.
A test is fully passed only if every run passed for that test.Wrong answer: 11Did not follow instructions: 1Invalid tool call: 1Response Time (avg)3.12sResponse Time (max)44.76sResponse Time (total)68.69sA test is fully passed only if every run passed for that test.…
Attempt pass rate
43.9%Attempt pass rate = passed attempts / total attempts across runs.…
40.9%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
0.208Shows the average cost per correct benchmark answer in cents (lower is better).…
N/AShows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$0.023Total Cost (Current Price)…
N/ATotal Cost (Current Price)…
Input Price
$0.120 / 1MInput Price…
N/AInput Price…
Output Price
$0.400 / 1MOutput Price…
N/AOutput Price…
Total Input Tokens
131,257Total Input Tokens…
122,463Total Input Tokens…
Output Tokens
15,781Output Tokens…
11,445Output Tokens…
Reasoning Tokens
0Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
7.58sResponse Time (avg)…
3.12sResponse Time (avg)…
Response Time (max)
57.10sResponse Time (max)…
44.76sResponse Time (max)…
Response Time (total)
166.82sResponse Time (total)…
68.69sResponse Time (total)…
Parameters
25.2B total (3.8B active)
27.3B
Availability
Open source
Weights available
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#204 Gemma 4 26B A4B
none
Cost
$0.001
Time
39.5s
Tokens
790 tok
#211 Qwen3.8 27B
noneCandidate responses were generated on local hardware. OpenRouter was used only for judging.
8.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
75.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.28sResponse Time (max)2.09sResponse Time (total)5.13sA test is fully passed only if every run passed for that test.…
1.28sResponse Time (avg)…
852Total Input Tokens…
230Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)828msResponse Time (max)1.95sResponse Time (total)3.31sA test is fully passed only if every run passed for that test.…
3.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Timed out: 1Response Time (avg)4.16sResponse Time (max)7.07sResponse Time (total)12.48sA test is fully passed only if every run passed for that test.…
4.16sResponse Time (avg)…
7,736Total Input Tokens…
476Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)1.19sResponse Time (max)1.97sResponse Time (total)3.56sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Wrong answer: 1Response Time (avg)37.23sResponse Time (max)43.93sResponse Time (total)74.46sA test is fully passed only if every run passed for that test.…
37.23sResponse Time (avg)…
104,894Total Input Tokens…
14,266Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Wrong answer: 1Response Time (avg)24.10sResponse Time (max)44.76sResponse Time (total)48.21sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.70sResponse Time (max)2.21sResponse Time (total)3.41sA test is fully passed only if every run passed for that test.…
1.70sResponse Time (avg)…
8,352Total Input Tokens…
285Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.14sResponse Time (max)1.18sResponse Time (total)2.27sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)2.10sResponse Time (max)4.23sResponse Time (total)6.31sA test is fully passed only if every run passed for that test.…
2.10sResponse Time (avg)…
878Total Input Tokens…
27Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)552msResponse Time (max)675msResponse Time (total)1.66sA test is fully passed only if every run passed for that test.…
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)3.54sResponse Time (max)3.54sResponse Time (total)3.54sA test is fully passed only if every run passed for that test.…
3.54sResponse Time (avg)…
576Total Input Tokens…
85Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
4.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.31sResponse Time (max)1.31sResponse Time (total)1.31sA test is fully passed only if every run passed for that test.…
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)690msResponse Time (max)878msResponse Time (total)1.38sA test is fully passed only if every run passed for that test.…
690msResponse Time (avg)…
795Total Input Tokens…
75Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)582msResponse Time (max)633msResponse Time (total)1.16sA test is fully passed only if every run passed for that test.…
6.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)744msResponse Time (max)972msResponse Time (total)2.23sA test is fully passed only if every run passed for that test.…
744msResponse Time (avg)…
828Total Input Tokens…
114Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)1.25sResponse Time (max)1.84sResponse Time (total)3.76sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)57.10sResponse Time (max)57.10sResponse Time (total)57.10sA test is fully passed only if every run passed for that test.…
57.10sResponse Time (avg)…
6,123Total Input Tokens…
210Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.75sResponse Time (max)2.75sResponse Time (total)2.75sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)778msResponse Time (max)778msResponse Time (total)778msA test is fully passed only if every run passed for that test.…
778msResponse Time (avg)…
223Total Input Tokens…
13Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)703msResponse Time (max)703msResponse Time (total)703msA test is fully passed only if every run passed for that test.…