The average score is effectively tied at 8.4 vs 8.4. Local hardware cost was not measured, so cost comparisons are unavailable. Grok 4.6 (medium) is faster at 73.46s vs 167.89s, with pass rates of 77.3% vs 81.8%.
Last updated at:
2026-08-15
Rank
#42
Total Output Tokens
656,635
Response Time (avg)
167.89s
Total Cost
N/A
Rank
#39
Total Output Tokens
249,808
Response Time (avg)
73.46s
Total Cost
$1.713
Recommended modelGrok 4.6 (medium)
It has the best score here (8.4), while responding about 2.3x faster than Qwen3.8 27B (high).
Detailed comparison
Metric
Qwen3.8 27BQwen3.8 27BhighCandidate responses were generated on local hardware. OpenRouter was used only for judging.Release: 2026-08-14
8.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#42
#39
Reliability
9.6First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
9.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
8.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
66/66
66/66
Tests Correct
A test is fully passed only if every run passed for that test.Did not follow instructions: 2No answer: 2Wrong answer: 2Response Time (avg)167.89sResponse Time (max)640.85sResponse Time (total)3693.54sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 5Did not follow instructions: 1Response Time (avg)73.46sResponse Time (max)514.35sResponse Time (total)1616.12sA test is fully passed only if every run passed for that test.…
Attempt pass rate
77.3%Attempt pass rate = passed attempts / total attempts across runs.…
81.8%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
4Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
N/AShows the average cost per correct benchmark answer in cents (lower is better).…
10.703Shows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
N/ATotal Cost (Current Price)…
$1.713Total Cost (Current Price)…
Input Price
N/AInput Price…
$2.000 / 1MInput Price…
Output Price
N/AOutput Price…
$6.000 / 1MOutput Price…
Total Input Tokens
100,179Total Input Tokens…
106,787Total Input Tokens…
Output Tokens
904Output Tokens…
4,811Output Tokens…
Reasoning Tokens
655,731Reasoning Tokens…
244,997Reasoning Tokens…
Response Time (avg)
167.89sResponse Time (avg)…
73.46sResponse Time (avg)…
Response Time (max)
640.85sResponse Time (max)…
514.35sResponse Time (max)…
Response Time (total)
3693.54sResponse Time (total)…
1616.12sResponse Time (total)…
Parameters
27.3B
~1.7T total (~170B active)
Availability
Weights available
Closed
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#42 Qwen3.8 27B
highCandidate responses were generated on local hardware. OpenRouter was used only for judging.
Cost
N/A
Time
270.1s
Tokens
18,963 tok
#39 SpaceXAI: Grok 4.6
medium
Cost
$0.028
Time
72.1s
Tokens
4,771 tok
Score
-
Cost
-
Time
-
Tokens
-
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
Anti-AI Tricks
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)8.10sResponse Time (max)21.48sResponse Time (total)32.40sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)10.84sResponse Time (max)16.03sResponse Time (total)43.37sA test is fully passed only if every run passed for that test.…
10.84sResponse Time (avg)…
2,991Total Input Tokens…
78Output Tokens…
6,335Reasoning Tokens…
Coding
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
8.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
88.9%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)248.62sResponse Time (max)458.25sResponse Time (total)745.85sA test is fully passed only if every run passed for that test.…
8.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
88.9%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)120.05sResponse Time (max)183.47sResponse Time (total)360.14sA test is fully passed only if every run passed for that test.…
120.05sResponse Time (avg)…
9,579Total Input Tokens…
351Output Tokens…
64,226Reasoning Tokens…
Combined
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)185.57sResponse Time (max)333.21sResponse Time (total)371.15sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)62.66sResponse Time (max)80.72sResponse Time (total)125.33sA test is fully passed only if every run passed for that test.…
62.66sResponse Time (avg)…
68,060Total Input Tokens…
3,687Output Tokens…
21,366Reasoning Tokens…
Data parsing and extraction
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)34.85sResponse Time (max)57.71sResponse Time (total)69.71sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)11.46sResponse Time (max)12.31sResponse Time (total)22.92sA test is fully passed only if every run passed for that test.…
11.46sResponse Time (avg)…
8,937Total Input Tokens…
225Output Tokens…
4,400Reasoning Tokens…
Domain specific
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Wrong answer: 1Response Time (avg)526.74sResponse Time (max)640.85sResponse Time (total)1580.21sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)304.33sResponse Time (max)514.35sResponse Time (total)912.98sA test is fully passed only if every run passed for that test.…
304.33sResponse Time (avg)…
2,523Total Input Tokens…
15Output Tokens…
127,702Reasoning Tokens…
General Intelligence
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)8.45sResponse Time (max)8.45sResponse Time (total)8.45sA test is fully passed only if every run passed for that test.…
4.1Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
2.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)17.27sResponse Time (max)17.27sResponse Time (total)17.27sA test is fully passed only if every run passed for that test.…
17.27sResponse Time (avg)…
1,077Total Input Tokens…
47Output Tokens…
2,052Reasoning Tokens…
Instructions following
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)8.24sResponse Time (max)10.18sResponse Time (total)16.49sA test is fully passed only if every run passed for that test.…
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)12.75sResponse Time (max)17.51sResponse Time (total)25.51sA test is fully passed only if every run passed for that test.…
12.75sResponse Time (avg)…
1,857Total Input Tokens…
57Output Tokens…
4,500Reasoning Tokens…
Puzzle Solving
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
7.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
77.8%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)199.90sResponse Time (max)569.33sResponse Time (total)599.69sA test is fully passed only if every run passed for that test.…
8.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
88.9%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)16.30sResponse Time (max)22.22sResponse Time (total)48.89sA test is fully passed only if every run passed for that test.…
16.30sResponse Time (avg)…
2,430Total Input Tokens…
123Output Tokens…
5,799Reasoning Tokens…
Tool Calling
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)5.62sResponse Time (max)5.62sResponse Time (total)5.62sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)12.83sResponse Time (max)12.83sResponse Time (total)12.83sA test is fully passed only if every run passed for that test.…
12.83sResponse Time (avg)…
8,544Total Input Tokens…
216Output Tokens…
1,068Reasoning Tokens…
Trivia
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Qwen3.8 27BCandidate responses were generated on local hardware. OpenRouter was used only for judging.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)263.98sResponse Time (max)263.98sResponse Time (total)263.98sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)46.88sResponse Time (max)46.88sResponse Time (total)46.88sA test is fully passed only if every run passed for that test.…