The average score is effectively tied at 7.2 vs 7.2. Qwen3.8 27B (medium) has the lower benchmark cost at ~$0.073 vs $0.323. Gemini 3.8 Flash (low) is faster at 7.54s vs 40.09s, with pass rates of 81.2% vs 66.7%.
Last updated at:
2026-10-01
Compared models
Rank
#142
Total Output Tokens
42,741
Response Time (avg)
7.54s
Total Cost
$0.323
Candidate responses were generated on local hardware. OpenRouter was used only for judging.
Rank
#141
Total Output Tokens
170,696
Response Time (avg)
40.09s
Total Cost
~$0.073Estimated GPU electricity cost only. It excludes the rest of the computer and hardware ownership. NVIDIA GeForce RTX 3090: 0.7349 h × 0.300 kW × €0.2896/kWh × 1.138 USD/EUR = ~$0.072660.
Recommended modelQwen3.8 27B (medium)
It has the best score here (7.2), while costing about 4.4x less than Gemini 3.8 Flash (low).
Qwen3.8 27BQwen3.8 27BmediumCandidate responses were generated on local hardware. OpenRouter was used only for judging.Release: 2026-08-14Free Available
Qwen3.8 27BQwen3.8 27BmediumCandidate responses were generated on local hardware. OpenRouter was used only for judging.Release: 2026-08-14Free Available
Score
7.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#142
#141
Reliability
9.9First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
9.7First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
8.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
9.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
69/69
69/69
Tests Correct
A test is fully passed only if every run passed for that test.Wrong answer: 6No answer: 1Response Time (avg)7.54sResponse Time (max)67.18sResponse Time (total)173.36sA test is fully passed only if every run passed for that test.…
6.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)67.18sResponse Time (max)67.18sResponse Time (total)67.18sA test is fully passed only if every run passed for that test.…
67.18sResponse Time (avg)…
104,063Total Input Tokens…
1,417Output Tokens…
0Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
4.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)187.85sResponse Time (max)187.85sResponse Time (total)187.85sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.96sResponse Time (max)2.77sResponse Time (total)7.83sA test is fully passed only if every run passed for that test.…
1.96sResponse Time (avg)…
492Total Input Tokens…
171Output Tokens…
1,054Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.93sResponse Time (max)5.18sResponse Time (total)11.70sA test is fully passed only if every run passed for that test.…
5.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)4.31sResponse Time (max)5.62sResponse Time (total)12.92sA test is fully passed only if every run passed for that test.…
4.31sResponse Time (avg)…
8,118Total Input Tokens…
449Output Tokens…
5,706Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)40.50sResponse Time (max)72.14sResponse Time (total)121.51sA test is fully passed only if every run passed for that test.…
3.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Wrong answer: 1Response Time (avg)19.61sResponse Time (max)35.14sResponse Time (total)39.21sA test is fully passed only if every run passed for that test.…
19.61sResponse Time (avg)…
87,701Total Input Tokens…
6,571Output Tokens…
14,001Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
8.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Response Time (avg)124.48sResponse Time (max)214.63sResponse Time (total)248.96sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.34sResponse Time (max)3.48sResponse Time (total)4.69sA test is fully passed only if every run passed for that test.…
2.34sResponse Time (avg)…
7,548Total Input Tokens…
279Output Tokens…
1,000Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)8.76sResponse Time (max)10.36sResponse Time (total)17.53sA test is fully passed only if every run passed for that test.…
5.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)6.50sResponse Time (max)9.32sResponse Time (total)19.51sA test is fully passed only if every run passed for that test.…
6.50sResponse Time (avg)…
642Total Input Tokens…
12Output Tokens…
7,051Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)76.98sResponse Time (max)144.07sResponse Time (total)230.93sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.89sResponse Time (max)1.89sResponse Time (total)1.89sA test is fully passed only if every run passed for that test.…
1.89sResponse Time (avg)…
486Total Input Tokens…
72Output Tokens…
316Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)9.20sResponse Time (max)9.20sResponse Time (total)9.20sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.79sResponse Time (max)3.11sResponse Time (total)5.57sA test is fully passed only if every run passed for that test.…
2.79sResponse Time (avg)…
615Total Input Tokens…
72Output Tokens…
1,873Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.54sResponse Time (max)2.67sResponse Time (total)5.07sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.25sResponse Time (max)3.00sResponse Time (total)6.75sA test is fully passed only if every run passed for that test.…
2.25sResponse Time (avg)…
558Total Input Tokens…
165Output Tokens…
1,458Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
8.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)12.62sResponse Time (max)29.62sResponse Time (total)37.85sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)5.51sResponse Time (max)5.51sResponse Time (total)5.51sA test is fully passed only if every run passed for that test.…
5.51sResponse Time (avg)…
5,457Total Input Tokens…
233Output Tokens…
224Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.29sResponse Time (max)2.29sResponse Time (total)2.29sA test is fully passed only if every run passed for that test.…
2.29sResponse Time (avg)…
156Total Input Tokens…
11Output Tokens…
606Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)11.29sResponse Time (max)11.29sResponse Time (total)11.29sA test is fully passed only if every run passed for that test.…