Qwen3.8 27B leads on average score with 5.5 vs 5.4. Local hardware cost was not measured, so cost comparisons are unavailable. Qwen3.8 27B is faster at 3.12s vs 10.63s, with pass rates of 43.9% vs 40.9%.
Last updated at:
2026-08-15
Rank
#215
Total Output Tokens
60,414
Response Time (avg)
10.63s
Total Cost
$0.046
Rank
#211
Total Output Tokens
11,445
Response Time (avg)
3.12s
Total Cost
N/A
Recommended modelQwen3.8 27B
It has the best score here (5.5), while responding about 3.4x faster than KAT-Coder-Air V2.5 (low).
A test is fully passed only if every run passed for that test.Wrong answer: 11Did not follow instructions: 1Invalid tool call: 1Response Time (avg)3.12sResponse Time (max)44.76sResponse Time (total)68.69sA test is fully passed only if every run passed for that test.…
Attempt pass rate
43.9%Attempt pass rate = passed attempts / total attempts across runs.…
40.9%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
6Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
0.649Shows the average cost per correct benchmark answer in cents (lower is better).…
N/AShows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$0.046Total Cost (Current Price)…
N/ATotal Cost (Current Price)…
Input Price
$0.150 / 1MInput Price…
N/AInput Price…
Output Price
$0.600 / 1MOutput Price…
N/AOutput Price…
Total Input Tokens
61,094Total Input Tokens…
122,463Total Input Tokens…
Output Tokens
6,532Output Tokens…
11,445Output Tokens…
Reasoning Tokens
53,882Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
10.63sResponse Time (avg)…
3.12sResponse Time (avg)…
Response Time (max)
86.23sResponse Time (max)…
44.76sResponse Time (max)…
Response Time (total)
233.77sResponse Time (total)…
68.69sResponse Time (total)…
Parameters
~35B total (~3B active)
27.3B
Availability
Closed
Weights available
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#215 KAT-Coder-Air V2.5
low
Cost
$0.004
Time
29.9s
Tokens
6,008 tok
#211 Qwen3.8 27B
noneCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)3.50sResponse Time (max)5.89sResponse Time (total)13.99sA test is fully passed only if every run passed for that test.…
3.50sResponse Time (avg)…
672Total Input Tokens…
422Output Tokens…
1,548Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)828msResponse Time (max)1.95sResponse Time (total)3.31sA test is fully passed only if every run passed for that test.…
3.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
11.1%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Extra formatting: 2Wrong answer: 1Response Time (avg)11.06sResponse Time (max)14.01sResponse Time (total)33.18sA test is fully passed only if every run passed for that test.…
11.06sResponse Time (avg)…
7,893Total Input Tokens…
2,001Output Tokens…
13,623Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)1.19sResponse Time (max)1.97sResponse Time (total)3.56sA test is fully passed only if every run passed for that test.…
6.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)55.86sResponse Time (max)86.23sResponse Time (total)111.72sA test is fully passed only if every run passed for that test.…
55.86sResponse Time (avg)…
33,753Total Input Tokens…
2,046Output Tokens…
17,833Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 1Wrong answer: 1Response Time (avg)24.10sResponse Time (max)44.76sResponse Time (total)48.21sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)2.82sResponse Time (max)3.26sResponse Time (total)5.64sA test is fully passed only if every run passed for that test.…
2.82sResponse Time (avg)…
7,776Total Input Tokens…
261Output Tokens…
1,069Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.14sResponse Time (max)1.18sResponse Time (total)2.27sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Extra formatting: 2API error: 1Response Time (avg)8.91sResponse Time (max)13.68sResponse Time (total)26.72sA test is fully passed only if every run passed for that test.…
8.91sResponse Time (avg)…
562Total Input Tokens…
937Output Tokens…
10,224Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)552msResponse Time (max)675msResponse Time (total)1.66sA test is fully passed only if every run passed for that test.…
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)10.11sResponse Time (max)10.11sResponse Time (total)10.11sA test is fully passed only if every run passed for that test.…
10.11sResponse Time (avg)…
516Total Input Tokens…
142Output Tokens…
596Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
4.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.31sResponse Time (max)1.31sResponse Time (total)1.31sA test is fully passed only if every run passed for that test.…
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.64sResponse Time (max)2.19sResponse Time (total)3.28sA test is fully passed only if every run passed for that test.…
1.64sResponse Time (avg)…
699Total Input Tokens…
89Output Tokens…
640Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)582msResponse Time (max)633msResponse Time (total)1.16sA test is fully passed only if every run passed for that test.…
3.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Did not follow instructions: 1Response Time (avg)1.57sResponse Time (max)1.86sResponse Time (total)4.71sA test is fully passed only if every run passed for that test.…
1.57sResponse Time (avg)…
696Total Input Tokens…
294Output Tokens…
796Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
6.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)1.25sResponse Time (max)1.84sResponse Time (total)3.76sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)4.47sResponse Time (max)4.47sResponse Time (total)4.47sA test is fully passed only if every run passed for that test.…
4.47sResponse Time (avg)…
8,391Total Input Tokens…
330Output Tokens…
1,065Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.75sResponse Time (max)2.75sResponse Time (total)2.75sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)19.98sResponse Time (max)19.98sResponse Time (total)19.98sA test is fully passed only if every run passed for that test.…
19.98sResponse Time (avg)…
136Total Input Tokens…
10Output Tokens…
6,488Reasoning Tokens…
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)703msResponse Time (max)703msResponse Time (total)703msA test is fully passed only if every run passed for that test.…