Gemini 3 PRO Preview (medium) leads on average score with 6.0 vs 5.6. GLM 5V Turbo has the lower benchmark cost at $0.052 vs $0.385. GLM 5V Turbo is faster at 2.99s vs 9.05s, with pass rates of 63.6% vs 36.4%.
Last updated at:
2026-09-22
Compared models
Rank
#225
Total Output Tokens
11,592
Response Time (avg)
9.05s
Total Cost
$0.385
Rank
#248
Total Output Tokens
1,766
Response Time (avg)
2.99s
Total Cost
$0.052
Recommended modelGemini 3 PRO Preview (medium)
It has the strongest score in this comparison (6.0) and the best overall balance of cost and response time across all 2 models.
GLM 5V TurboGLM 5V TurbononeArchived model: this model is no longer updated or tested on new tests.Release: 2026-04-01
Score
6.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#225
#248
Reliability
N/AFirst-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
9.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
9.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
63/66
63/66
Tests Correct
A test is fully passed only if every run passed for that test.API error: 4Wrong answer: 3Response Time (avg)9.05sResponse Time (max)26.24sResponse Time (total)90.53sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 11Did not follow instructions: 2Response Time (avg)2.99sResponse Time (max)6.51sResponse Time (total)62.74sA test is fully passed only if every run passed for that test.…
Attempt pass rate
63.6%Attempt pass rate = passed attempts / total attempts across runs.…
36.4%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
63Total Runs…
63Total Runs…
Cost per result
1.406
0.645
Total Cost
$0.385
$0.052
Input Price
$9.506 / 1MInput Price…
$1.200 / 1MInput Price…
Output Price
$9.506 / 1MOutput Price…
$4.000 / 1MOutput Price…
Total Input Tokens
28,848Total Input Tokens…
37,100Total Input Tokens…
Output Tokens
1,490Output Tokens…
1,766Output Tokens…
Reasoning Tokens
10,102Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
9.05sResponse Time (avg)…
2.99sResponse Time (avg)…
Response Time (max)
26.24sResponse Time (max)…
6.51sResponse Time (max)…
Response Time (total)
90.53sResponse Time (total)…
62.74sResponse Time (total)…
Parameters
~1.2T total (~20B active)
~106B total (~12B active)
Availability
Closed
Closed
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#225 Gemini 3 PRO Preview
medium
No endpoints found for google/gemini-3-pro-preview.
Cost
$0.000
Time
0.1s
Tokens
0 tok
#248 GLM 5V Turbo
none
Cost
$0.042
Time
177.3s
Tokens
10,434 tok
Rank
-
Cost
-
Time
-
Tokens
-
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
Anti-AI Tricks
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)14.99sResponse Time (max)26.24sResponse Time (total)29.99sA test is fully passed only if every run passed for that test.…
14.99sResponse Time (avg)…
500Total Input Tokens…
149Output Tokens…
1,485Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
4.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)3.13sResponse Time (max)5.90sResponse Time (total)12.50sA test is fully passed only if every run passed for that test.…
3.13sResponse Time (avg)…
555Total Input Tokens…
281Output Tokens…
0Reasoning Tokens…
Coding
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0msResponse Time (avg)…
0Total Input Tokens…
0Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
5.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)3.13sResponse Time (max)5.30sResponse Time (total)9.40sA test is fully passed only if every run passed for that test.…
3.13sResponse Time (avg)…
7,256Total Input Tokens…
360Output Tokens…
0Reasoning Tokens…
Combined
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)10.37sResponse Time (max)10.37sResponse Time (total)10.37sA test is fully passed only if every run passed for that test.…
10.37sResponse Time (avg)…
13,211Total Input Tokens…
351Output Tokens…
952Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)6.51sResponse Time (max)6.51sResponse Time (total)6.51sA test is fully passed only if every run passed for that test.…
6.51sResponse Time (avg)…
12,708Total Input Tokens…
276Output Tokens…
0Reasoning Tokens…
Data parsing and extraction
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)10.84sResponse Time (max)10.84sResponse Time (total)10.84sA test is fully passed only if every run passed for that test.…
10.84sResponse Time (avg)…
7,259Total Input Tokens…
279Output Tokens…
3,156Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.81sResponse Time (max)5.69sResponse Time (total)7.62sA test is fully passed only if every run passed for that test.…
3.81sResponse Time (avg)…
7,107Total Input Tokens…
204Output Tokens…
0Reasoning Tokens…
Domain specific
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)7.01sResponse Time (max)7.01sResponse Time (total)7.01sA test is fully passed only if every run passed for that test.…
7.01sResponse Time (avg)…
643Total Input Tokens…
15Output Tokens…
1,195Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)2.09sResponse Time (max)2.39sResponse Time (total)6.26sA test is fully passed only if every run passed for that test.…
2.09sResponse Time (avg)…
687Total Input Tokens…
24Output Tokens…
0Reasoning Tokens…
General Intelligence
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)9.34sResponse Time (max)9.34sResponse Time (total)9.34sA test is fully passed only if every run passed for that test.…
9.34sResponse Time (avg)…
486Total Input Tokens…
78Output Tokens…
374Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
4.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)2.22sResponse Time (max)2.22sResponse Time (total)2.22sA test is fully passed only if every run passed for that test.…
2.22sResponse Time (avg)…
477Total Input Tokens…
114Output Tokens…
0Reasoning Tokens…
Instructions following
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
9.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.26sResponse Time (max)3.26sResponse Time (total)3.26sA test is fully passed only if every run passed for that test.…
3.26sResponse Time (avg)…
623Total Input Tokens…
69Output Tokens…
754Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.97sResponse Time (max)2.43sResponse Time (total)3.93sA test is fully passed only if every run passed for that test.…
1.97sResponse Time (avg)…
636Total Input Tokens…
60Output Tokens…
0Reasoning Tokens…
Puzzle Solving
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.88sResponse Time (max)4.23sResponse Time (total)7.77sA test is fully passed only if every run passed for that test.…
3.88sResponse Time (avg)…
570Total Input Tokens…
225Output Tokens…
1,215Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)2.40sResponse Time (max)3.81sResponse Time (total)7.21sA test is fully passed only if every run passed for that test.…
2.40sResponse Time (avg)…
609Total Input Tokens…
210Output Tokens…
0Reasoning Tokens…
Tool Calling
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)11.96sResponse Time (max)11.96sResponse Time (total)11.96sA test is fully passed only if every run passed for that test.…
11.96sResponse Time (avg)…
5,556Total Input Tokens…
324Output Tokens…
971Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)4.86sResponse Time (max)4.86sResponse Time (total)4.86sA test is fully passed only if every run passed for that test.…
4.86sResponse Time (avg)…
6,879Total Input Tokens…
222Output Tokens…
0Reasoning Tokens…
Trivia
Score
Consistency
Attempt pass rate
Flaky tests
Tests Correct
Response Time (avg)
Input Tokens
Output Tokens
Reasoning Tokens
Gemini 3 PRO PreviewArchived model: this model is no longer updated or tested on new tests.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0msResponse Time (avg)…
0Total Input Tokens…
0Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)2.23sResponse Time (max)2.23sResponse Time (total)2.23sA test is fully passed only if every run passed for that test.…