Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Gemma 4 26B A4B (medium) vs GLM 5.2

The average score is effectively tied at 6.6 vs 6.6. Gemma 4 26B A4B (medium) has the lower benchmark cost at $0.089 vs $0.120. GLM 5.2 is faster at 9.34s vs 103.83s, with pass rates of 66.7% vs 59.1%.

Last updated at: 2026-08-01

Rank
#113
Total Output Tokens
247,527
Response Time (avg)
103.83s
Total Cost
$0.089
Rank
#114
Total Output Tokens
14,340
Response Time (avg)
9.34s
Total Cost
$0.120
Recommended model GLM 5.2

It has the best score here (6.6), while responding about 11.1x faster than Gemma 4 26B A4B (medium).

Detailed comparison

Metric Gemma 4 26B A4B Gemma 4 26B A4B medium Release: 2026-04-03 Free Available GLM 5.2 GLM 5.2 none Release: 2026-06-17
Score 6.6 6.6
Rank #113 #114
Reliability 9.4 10.0
Consistency 9.2 9.2
Tests Correct
Attempt pass rate 66.7% 59.1%
Flaky tests 2 2
Total Runs 66 64
Cost per result 0.643 1.421
Total Cost $0.089 $0.120
Input Price $0.070 / 1M $0.761 / 1M
Output Price $0.340 / 1M $2.389 / 1M
Total Input Tokens 77,550 112,359
Output Tokens 28,036 14,340
Reasoning Tokens 219,491 0
Response Time (avg) 103.83s 9.34s
Response Time (max) 912.19s 79.65s
Response Time (total) 2180.47s 205.46s

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#113 Gemma 4 26B A4B

medium
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

#114 GLM 5.2

none
Invalid SVG
Cost
$0.033
Time
87.7s
Tokens
7,455 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemma 4 26B A4B 2.9 10.0 0.0% 0 272.54s 5,062 14,838 44,567
GLM 5.2 3.7 9.5 0.0% 0 7.55s 7,263 1,958 0

Quick Compare

Switch Comparison Pair