Qwen3.5-122B-A10B vs GLM 5V Turbo benchmark comparison: GLM 5V Turbo leads on average score with 5.9 vs 5.3. Qwen3.5-122B-A10B has the lower benchmark cost at $0.020 vs $0.052. GLM 5V Turbo is faster at 2.99s vs 3.41s, with pass rates of 31.8% vs 38.1%.
Recommended model: Qwen3.5-122B-A10B - Its score stays close to the best score here (5.3 vs 5.9), while costing about 2.7x less than GLM 5V Turbo.
GLM 5V TurboGLM 5V TurbononeArchived model: this model is no longer updated or tested on new tests.Release: 2026-04-01
Score
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#125
#105
Reliability
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
9.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Tests Correct
A test is fully passed only if every run passed for that test.Wrong answer: 13Did not follow instructions: 2Response Time (avg)3.41sResponse Time (max)46.00sResponse Time (total)71.59sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 11Did not follow instructions: 2Response Time (avg)2.99sResponse Time (max)6.51sResponse Time (total)62.74sA test is fully passed only if every run passed for that test.…
Attempt pass rate
31.8%Attempt pass rate = passed attempts / total attempts across runs.…
38.1%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
63Total Runs…
63Total Runs…
Cost per result
0.393Shows the average cost per correct benchmark answer in cents (lower is better).…
0.645Shows the average cost per correct benchmark answer in cents (lower is better).…
Total Cost
$0.020Total Cost (Current Price)…
$0.052Total Cost (Current Price)…
Input Price
$0.260 / 1MInput Price…
$1.200 / 1MInput Price…
Output Price
$2.080 / 1MOutput Price…
$4.000 / 1MOutput Price…
Total Input Tokens
47,735Total Input Tokens…
37,100Total Input Tokens…
Output Tokens
3,383Output Tokens…
1,766Output Tokens…
Reasoning Tokens
0Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
3.41sResponse Time (avg)…
2.99sResponse Time (avg)…
Response Time (max)
46.00sResponse Time (max)…
6.51sResponse Time (max)…
Response Time (total)
71.59sResponse Time (total)…
62.74sResponse Time (total)…
Generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
4.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)1.59sResponse Time (max)3.60sResponse Time (total)6.38sA test is fully passed only if every run passed for that test.…
1.59sResponse Time (avg)…
696Total Input Tokens…
312Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
4.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)3.13sResponse Time (max)5.90sResponse Time (total)12.50sA test is fully passed only if every run passed for that test.…
3.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)2.77sResponse Time (max)4.03sResponse Time (total)8.32sA test is fully passed only if every run passed for that test.…
2.77sResponse Time (avg)…
7,913Total Input Tokens…
693Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
5.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)3.13sResponse Time (max)5.30sResponse Time (total)9.40sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)46.00sResponse Time (max)46.00sResponse Time (total)46.00sA test is fully passed only if every run passed for that test.…
46.00sResponse Time (avg)…
20,175Total Input Tokens…
1,137Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)6.51sResponse Time (max)6.51sResponse Time (total)6.51sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)1.01sResponse Time (max)1.06sResponse Time (total)2.02sA test is fully passed only if every run passed for that test.…
1.01sResponse Time (avg)…
7,794Total Input Tokens…
243Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)3.81sResponse Time (max)5.69sResponse Time (total)7.62sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)465msResponse Time (max)492msResponse Time (total)1.39sA test is fully passed only if every run passed for that test.…
465msResponse Time (avg)…
789Total Input Tokens…
15Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)2.09sResponse Time (max)2.39sResponse Time (total)6.26sA test is fully passed only if every run passed for that test.…
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)1.12sResponse Time (max)1.12sResponse Time (total)1.12sA test is fully passed only if every run passed for that test.…
1.12sResponse Time (avg)…
522Total Input Tokens…
66Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
4.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Response Time (avg)2.22sResponse Time (max)2.22sResponse Time (total)2.22sA test is fully passed only if every run passed for that test.…
6.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)513msResponse Time (max)570msResponse Time (total)1.03sA test is fully passed only if every run passed for that test.…
513msResponse Time (avg)…
711Total Input Tokens…
69Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)1.97sResponse Time (max)2.43sResponse Time (total)3.93sA test is fully passed only if every run passed for that test.…
3.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Did not follow instructions: 1Response Time (avg)1.00sResponse Time (max)1.41sResponse Time (total)3.00sA test is fully passed only if every run passed for that test.…
1.00sResponse Time (avg)…
714Total Input Tokens…
575Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Did not follow instructions: 1Wrong answer: 1Response Time (avg)2.40sResponse Time (max)3.81sResponse Time (total)7.21sA test is fully passed only if every run passed for that test.…
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)2.04sResponse Time (max)2.04sResponse Time (total)2.04sA test is fully passed only if every run passed for that test.…
2.04sResponse Time (avg)…
8,211Total Input Tokens…
264Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
10.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
100.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)4.86sResponse Time (max)4.86sResponse Time (total)4.86sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)295msResponse Time (max)295msResponse Time (total)295msA test is fully passed only if every run passed for that test.…
295msResponse Time (avg)…
210Total Input Tokens…
9Output Tokens…
0Reasoning Tokens…
GLM 5V TurboArchived model: this model is no longer updated or tested on new tests.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)2.23sResponse Time (max)2.23sResponse Time (total)2.23sA test is fully passed only if every run passed for that test.…