7.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#48
#16
Reliability
9.8First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
N/AFirst-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
8.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Tests Correct
A test is fully passed only if every run passed for that test.Did not follow instructions: 3Wrong answer: 3No answer: 1Response Time (avg)61.96sResponse Time (max)149.23sResponse Time (total)1115.31sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)15.25sResponse Time (max)43.55sResponse Time (total)182.96sA test is fully passed only if every run passed for that test.…
Attempt pass rate
74.1%Attempt pass rate = passed attempts / total attempts across runs.…
75.0%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
4Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
54Total Runs…
57Total Runs…
Cost per result
18.579Shows the average cost per correct benchmark answer in cents (lower is better).…
0.000Shows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$2.044Total Cost…
$0.000Total Cost…
Input Price
$0.250 / 1MInput Price…
$0.000 / 1MInput Price…
Output Price
$1.500 / 1MOutput Price…
$0.000 / 1MOutput Price…
Output Tokens
1,984Output Tokens…
1,153Output Tokens…
Reasoning Tokens
1,355,583Reasoning Tokens…
62,197Reasoning Tokens…
Response Time (avg)
61.96sResponse Time (avg)…
15.25sResponse Time (avg)…
Response Time (max)
149.23sResponse Time (max)…
43.55sResponse Time (max)…
Response Time (total)
1115.31sResponse Time (total)…
182.96sResponse Time (total)…
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
Anti-AI Tricks
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
9.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)37.16sResponse Time (max)140.53sResponse Time (total)148.65sA test is fully passed only if every run passed for that test.…
37.16sResponse Time (avg)…
100Output Tokens…
130,598Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)11.69sResponse Time (max)19.37sResponse Time (total)35.08sA test is fully passed only if every run passed for that test.…
11.69sResponse Time (avg)…
61Output Tokens…
5,812Reasoning Tokens…
Coding
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)137.63sResponse Time (max)137.63sResponse Time (total)137.63sA test is fully passed only if every run passed for that test.…
137.63sResponse Time (avg)…
666Output Tokens…
188,733Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0msResponse Time (avg)…
0Output Tokens…
0Reasoning Tokens…
Combined
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)149.23sResponse Time (max)149.23sResponse Time (total)149.23sA test is fully passed only if every run passed for that test.…
149.23sResponse Time (avg)…
327Output Tokens…
198,243Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)34.95sResponse Time (max)34.95sResponse Time (total)34.95sA test is fully passed only if every run passed for that test.…
34.95sResponse Time (avg)…
452Output Tokens…
13,073Reasoning Tokens…
Data parsing and extraction
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)4.49sResponse Time (max)4.96sResponse Time (total)8.98sA test is fully passed only if every run passed for that test.…
4.49sResponse Time (avg)…
279Output Tokens…
7,351Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)14.95sResponse Time (max)15.40sResponse Time (total)29.90sA test is fully passed only if every run passed for that test.…
14.95sResponse Time (avg)…
270Output Tokens…
10,706Reasoning Tokens…
Domain specific
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
3.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)139.90sResponse Time (max)141.40sResponse Time (total)419.69sA test is fully passed only if every run passed for that test.…
139.90sResponse Time (avg)…
18Output Tokens…
566,210Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)22.08sResponse Time (max)43.55sResponse Time (total)66.23sA test is fully passed only if every run passed for that test.…
22.08sResponse Time (avg)…
49Output Tokens…
26,895Reasoning Tokens…
General Intelligence
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
2.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)45.69sResponse Time (max)45.69sResponse Time (total)45.69sA test is fully passed only if every run passed for that test.…
45.69sResponse Time (avg)…
95Output Tokens…
64,644Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0msResponse Time (avg)…
0Output Tokens…
0Reasoning Tokens…
Instructions following
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
7.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Response Time (avg)23.26sResponse Time (max)43.87sResponse Time (total)46.51sA test is fully passed only if every run passed for that test.…
23.26sResponse Time (avg)…
52Output Tokens…
3,549Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.40sResponse Time (max)3.40sResponse Time (total)3.40sA test is fully passed only if every run passed for that test.…
3.40sResponse Time (avg)…
27Output Tokens…
1,383Reasoning Tokens…
Puzzle Solving
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
5.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 2Response Time (avg)50.83sResponse Time (max)144.85sResponse Time (total)152.49sA test is fully passed only if every run passed for that test.…
50.83sResponse Time (avg)…
213Output Tokens…
193,654Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)7.52sResponse Time (max)7.52sResponse Time (total)7.52sA test is fully passed only if every run passed for that test.…
7.52sResponse Time (avg)…
27Output Tokens…
2,998Reasoning Tokens…
Tool Calling
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)6.44sResponse Time (max)6.44sResponse Time (total)6.44sA test is fully passed only if every run passed for that test.…
6.44sResponse Time (avg)…
234Output Tokens…
2,601Reasoning Tokens…
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)5.87sResponse Time (max)5.87sResponse Time (total)5.87sA test is fully passed only if every run passed for that test.…
5.87sResponse Time (avg)…
267Output Tokens…
1,330Reasoning Tokens…
Trivia
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Gemini 3.1 Flash LiteArchived model: this model is no longer updated or tested on new tests.
-
-
-
-
-
-
-
-
Qwen3.6 Plus PreviewArchived model: this model is no longer updated or tested on new tests.
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…