6.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#30
#12
Tests Correct
A test is fully passed only if every run passed for that test.Extra formatting: 4Wrong answer: 2Response Time (avg)25.08sResponse Time (max)83.40sResponse Time (total)200.67sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 4Response Time (avg)3.49sResponse Time (max)11.91sResponse Time (total)52.29sA test is fully passed only if every run passed for that test.…
Consistency
8.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Cost per result
14.411Shows the average cost per correct benchmark answer in cents (lower is better).…
0.170Shows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$1.297Total Cost…
$0.019Total Cost…
Attempt pass rate
64.4%Attempt pass rate = passed attempts / total attempts across runs.…
73.3%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
common.totalRuns
45 (15 x 3)common.totalRuns…
45 (15 x 3)common.totalRuns…
Output Tokens
26,066Output Tokens…
1,542Output Tokens…
Reasoning Tokens
17,071Reasoning Tokens…
6,888Reasoning Tokens…
Response Time (avg)
25.08sResponse Time (avg)…
3.49sResponse Time (avg)…
Response Time (max)
83.40sResponse Time (max)…
11.91sResponse Time (max)…
Response Time (total)
200.67sResponse Time (total)…
52.29sResponse Time (total)…
Top Models by Score
Score vs Total Cost
Response Time (avg)
Avg Score vs Response Time (avg)
Category Breakdown
Anti-AI Tricks
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
55.6%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Extra formatting: 2Response Time (avg)11.88sResponse Time (max)11.88sResponse Time (total)11.88sA test is fully passed only if every run passed for that test.…
11.88sResponse Time (avg)…
897Output Tokens…
1,000Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
7.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)2.18sResponse Time (max)3.18sResponse Time (total)6.53sA test is fully passed only if every run passed for that test.…
2.18sResponse Time (avg)…
456Output Tokens…
1,224Reasoning Tokens…
Combined
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)76.66sResponse Time (max)76.66sResponse Time (total)76.66sA test is fully passed only if every run passed for that test.…
76.66sResponse Time (avg)…
8,178Output Tokens…
5,194Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)11.91sResponse Time (max)11.91sResponse Time (total)11.91sA test is fully passed only if every run passed for that test.…
11.91sResponse Time (avg)…
225Output Tokens…
762Reasoning Tokens…
Data parsing and extraction
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
9.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)7.37sResponse Time (max)7.37sResponse Time (total)7.37sA test is fully passed only if every run passed for that test.…
7.37sResponse Time (avg)…
691Output Tokens…
757Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
9.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.00sResponse Time (max)3.74sResponse Time (total)5.99sA test is fully passed only if every run passed for that test.…
3.00sResponse Time (avg)…
291Output Tokens…
696Reasoning Tokens…
Domain specific
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Extra formatting: 2Wrong answer: 1Response Time (avg)83.40sResponse Time (max)83.40sResponse Time (total)83.40sA test is fully passed only if every run passed for that test.…
83.40sResponse Time (avg)…
14,642Output Tokens…
8,687Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)2.36sResponse Time (max)3.51sResponse Time (total)7.07sA test is fully passed only if every run passed for that test.…
2.36sResponse Time (avg)…
18Output Tokens…
1,212Reasoning Tokens…
Instructions following
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.43sResponse Time (max)2.43sResponse Time (total)2.43sA test is fully passed only if every run passed for that test.…
2.43sResponse Time (avg)…
266Output Tokens…
467Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.49sResponse Time (max)1.66sResponse Time (total)2.99sA test is fully passed only if every run passed for that test.…
1.49sResponse Time (avg)…
72Output Tokens…
753Reasoning Tokens…
Puzzle Solving
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
7.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)4.60sResponse Time (max)4.66sResponse Time (total)9.20sA test is fully passed only if every run passed for that test.…
4.60sResponse Time (avg)…
531Output Tokens…
637Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.76sResponse Time (max)5.08sResponse Time (total)8.27sA test is fully passed only if every run passed for that test.…
2.76sResponse Time (avg)…
243Output Tokens…
1,248Reasoning Tokens…
Tool Calling
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Output Tokens
Reasoning Tokens
Anthropic: Claude Opus 4.6
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)9.73sResponse Time (max)9.73sResponse Time (total)9.73sA test is fully passed only if every run passed for that test.…
9.73sResponse Time (avg)…
861Output Tokens…
329Reasoning Tokens…
Google: Gemini 3.1 Flash Lite Preview
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)9.54sResponse Time (max)9.54sResponse Time (total)9.54sA test is fully passed only if every run passed for that test.…