Gemini 3.5 Flash Lite (medium) vs Qwen3.8 27B (low)
Qwen3.8 27B (low) leads on average score with 7.9 vs 7.8. Local hardware cost was not measured, so cost comparisons are unavailable. Gemini 3.5 Flash Lite (medium) is faster at 5.12s vs 39.11s, with pass rates of 78.8% vs 68.2%.
Last updated at:
2026-08-15
Rank
#64
Total Output Tokens
96,235
Response Time (avg)
5.12s
Total Cost
$0.272
Rank
#61
Total Output Tokens
169,329
Response Time (avg)
39.11s
Total Cost
N/A
Recommended modelGemini 3.5 Flash Lite (medium)
Its score stays close to the best score here (7.8 vs 7.9), while responding about 7.6x faster than Qwen3.8 27B (low).
Qwen3.8 27BQwen3.8 27BlowCandidate responses were generated on local hardware. OpenRouter was used only for judging.Release: 2026-08-14
Score
7.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#64
#61
Reliability
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
7.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
66/66
66/66
Tests Correct
A test is fully passed only if every run passed for that test.Wrong answer: 6Did not follow instructions: 1Invalid tool call: 1Response Time (avg)5.12sResponse Time (max)26.41sResponse Time (total)112.73sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 4No answer: 2Did not follow instructions: 1Response Time (avg)39.11sResponse Time (max)376.04sResponse Time (total)860.51sA test is fully passed only if every run passed for that test.…
Attempt pass rate
78.8%Attempt pass rate = passed attempts / total attempts across runs.…
68.2%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
6Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
1.942Shows the average cost per correct benchmark answer in cents (lower is better).…
N/AShows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$0.272Total Cost (Current Price)…
N/ATotal Cost (Current Price)…
Input Price
$0.300 / 1MInput Price…
N/AInput Price…
Output Price
$2.500 / 1MOutput Price…
N/AOutput Price…
Total Input Tokens
103,916Total Input Tokens…
99,705Total Input Tokens…
Output Tokens
8,892Output Tokens…
1,545Output Tokens…
Reasoning Tokens
87,343Reasoning Tokens…
167,784Reasoning Tokens…
Response Time (avg)
5.12sResponse Time (avg)…
39.11sResponse Time (avg)…
Response Time (max)
26.41sResponse Time (max)…
376.04sResponse Time (max)…
Response Time (total)
112.73sResponse Time (total)…
860.51sResponse Time (total)…
Parameters
~150B total (~10B active)
27.3B
Availability
Closed
Weights available
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#64 Gemini 3.5 Flash Lite
medium
Cost
$0.010
Time
17.8s
Tokens
4,000 tok
#61 Qwen3.8 27B
lowCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.11sResponse Time (max)2.84sResponse Time (total)8.43sA test is fully passed only if every run passed for that test.…
2.11sResponse Time (avg)…
504Total Input Tokens…
150Output Tokens…
5,265Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.02sResponse Time (max)5.68sResponse Time (total)12.07sA test is fully passed only if every run passed for that test.…
7.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)9.96sResponse Time (max)15.08sResponse Time (total)29.87sA test is fully passed only if every run passed for that test.…
9.96sResponse Time (avg)…
8,122Total Input Tokens…
471Output Tokens…
28,930Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)140.92sResponse Time (max)376.04sResponse Time (total)422.75sA test is fully passed only if every run passed for that test.…
7.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Response Time (avg)17.87sResponse Time (max)26.41sResponse Time (total)35.73sA test is fully passed only if every run passed for that test.…
17.87sResponse Time (avg)…
79,299Total Input Tokens…
7,303Output Tokens…
29,209Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)90.63sResponse Time (max)143.71sResponse Time (total)181.27sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.11sResponse Time (max)2.63sResponse Time (total)4.21sA test is fully passed only if every run passed for that test.…
2.11sResponse Time (avg)…
7,362Total Input Tokens…
287Output Tokens…
2,226Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)8.03sResponse Time (max)9.09sResponse Time (total)16.06sA test is fully passed only if every run passed for that test.…
4.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.5%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)5.34sResponse Time (max)7.83sResponse Time (total)16.02sA test is fully passed only if every run passed for that test.…
5.34sResponse Time (avg)…
656Total Input Tokens…
17Output Tokens…
10,964Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)60.78sResponse Time (max)120.94sResponse Time (total)182.34sA test is fully passed only if every run passed for that test.…
4.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)2.97sResponse Time (max)2.97sResponse Time (total)2.97sA test is fully passed only if every run passed for that test.…
2.97sResponse Time (avg)…
488Total Input Tokens…
100Output Tokens…
1,558Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)5.37sResponse Time (max)5.37sResponse Time (total)5.37sA test is fully passed only if every run passed for that test.…
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.72sResponse Time (max)1.78sResponse Time (total)3.43sA test is fully passed only if every run passed for that test.…
1.72sResponse Time (avg)…
625Total Input Tokens…
69Output Tokens…
2,657Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.43sResponse Time (max)2.87sResponse Time (total)4.86sA test is fully passed only if every run passed for that test.…
8.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
88.9%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.82sResponse Time (max)2.25sResponse Time (total)5.47sA test is fully passed only if every run passed for that test.…
1.82sResponse Time (avg)…
570Total Input Tokens…
199Output Tokens…
2,988Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)4.91sResponse Time (max)8.81sResponse Time (total)14.74sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.84sResponse Time (max)2.84sResponse Time (total)2.84sA test is fully passed only if every run passed for that test.…
2.84sResponse Time (avg)…
6,132Total Input Tokens…
287Output Tokens…
1,277Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)8.42sResponse Time (max)8.42sResponse Time (total)8.42sA test is fully passed only if every run passed for that test.…
2.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)3.75sResponse Time (max)3.75sResponse Time (total)3.75sA test is fully passed only if every run passed for that test.…
3.75sResponse Time (avg)…
158Total Input Tokens…
9Output Tokens…
2,269Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)12.64sResponse Time (max)12.64sResponse Time (total)12.64sA test is fully passed only if every run passed for that test.…