Trinity Large Thinking (low) vs Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite leads on average score with 6.1 vs 6.0. Gemini 3.1 Flash Lite has the lower benchmark cost at $0.046 vs $0.656. Gemini 3.1 Flash Lite is faster at 1.75s vs 96.78s, with pass rates of 48.5% vs 50.0%.
Last updated at:
2026-07-28
Rank
#145
Total Output Tokens
815,220
Response Time (avg)
96.78s
Total Cost
$0.656
Rank
#137
Total Output Tokens
10,723
Response Time (avg)
1.75s
Total Cost
$0.046
Recommended modelGemini 3.1 Flash Lite
It has the best score here (6.1), while costing about 14.4x less than Trinity Large Thinking (low).
A test is fully passed only if every run passed for that test.Wrong answer: 11Did not follow instructions: 1No answer: 1Response Time (avg)1.75sResponse Time (max)16.25sResponse Time (total)38.60sA test is fully passed only if every run passed for that test.…
Attempt pass rate
48.5%Attempt pass rate = passed attempts / total attempts across runs.…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
6Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
4Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
66Total Runs…
Cost per result
8.199Shows the average cost per correct benchmark answer in cents (lower is better).…
0.507Shows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$0.656Total Cost (Current Price)…
$0.046Total Cost (Current Price)…
Input Price
$0.220 / 1MInput Price…
$0.250 / 1MInput Price…
Output Price
$0.850 / 1MOutput Price…
$1.500 / 1MOutput Price…
Total Input Tokens
122,845Total Input Tokens…
118,050Total Input Tokens…
Output Tokens
117,704Output Tokens…
10,723Output Tokens…
Reasoning Tokens
697,516Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
96.78sResponse Time (avg)…
1.75sResponse Time (avg)…
Response Time (max)
540.96sResponse Time (max)…
16.25sResponse Time (max)…
Response Time (total)
2129.13sResponse Time (total)…
38.60sResponse Time (total)…
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
8.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
75.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)5.70sResponse Time (max)10.16sResponse Time (total)22.82sA test is fully passed only if every run passed for that test.…
7.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
8.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)1.07sResponse Time (max)1.91sResponse Time (total)4.27sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)344.45sResponse Time (max)531.83sResponse Time (total)1033.35sA test is fully passed only if every run passed for that test.…
5.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)938msResponse Time (max)1.59sResponse Time (total)2.81sA test is fully passed only if every run passed for that test.…
5.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Invalid tool call: 2Response Time (avg)291.02sResponse Time (max)540.96sResponse Time (total)582.05sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No answer: 1Wrong answer: 1Response Time (avg)9.49sResponse Time (max)16.25sResponse Time (total)18.98sA test is fully passed only if every run passed for that test.…
7.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
83.3%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)14.30sResponse Time (max)21.07sResponse Time (total)28.60sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)843msResponse Time (max)907msResponse Time (total)1.69sA test is fully passed only if every run passed for that test.…
2.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.4Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2No answer: 1Response Time (avg)141.46sResponse Time (max)307.46sResponse Time (total)424.39sA test is fully passed only if every run passed for that test.…
2.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
11.1%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)762msResponse Time (max)814msResponse Time (total)2.29sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)12.65sResponse Time (max)12.65sResponse Time (total)12.65sA test is fully passed only if every run passed for that test.…
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)992msResponse Time (max)992msResponse Time (total)992msA test is fully passed only if every run passed for that test.…
4.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
16.7%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)4.07sResponse Time (max)4.47sResponse Time (total)8.14sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)859msResponse Time (max)975msResponse Time (total)1.72sA test is fully passed only if every run passed for that test.…
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)3.28sResponse Time (max)6.54sResponse Time (total)9.85sA test is fully passed only if every run passed for that test.…
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
4.8Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)720msResponse Time (max)826msResponse Time (total)2.16sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)5.19sResponse Time (max)5.19sResponse Time (total)5.19sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.97sResponse Time (max)2.97sResponse Time (total)2.97sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)2.10sResponse Time (max)2.10sResponse Time (total)2.10sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)733msResponse Time (max)733msResponse Time (total)733msA test is fully passed only if every run passed for that test.…