Navigate
AI BENCHY
Advertise here

Laguna XS 2.1 (medium) vs GLM 5.2

GLM 5.2 leads on average score with 6.7 vs 6.5. Laguna XS 2.1 (medium) has the lower benchmark cost at $0.068 vs $0.153. GLM 5.2 is faster at 9.65s vs 48.18s, with pass rates of 45.5% vs 63.6%.

Last updated at: 2026-08-20

Rank
#153
Total Output Tokens
522,480
Response Time (avg)
48.18s
Total Cost
$0.068
Rank
#138
Total Output Tokens
14,344
Response Time (avg)
9.65s
Total Cost
$0.153
Recommended model GLM 5.2

It has the best score here (6.7), while responding about 5.0x faster than Laguna XS 2.1 (medium).

Detailed comparison

Metric Laguna XS 2.1 Laguna XS 2.1 medium Release: 2026-07-02 Free Available GLM 5.2 GLM 5.2 none Release: 2026-06-17
Score 6.5 6.7
Rank #153 #138
Reliability 10.0 10.0
Consistency 9.2 9.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 45.5% 63.6%
Flaky tests 2 2
Total Runs 66 66
Cost per result 0.745 1.312
Total Cost $0.068 $0.153
Input Price $0.060 / 1M $0.966 / 1M
Output Price $0.120 / 1M $3.036 / 1M
Total Input Tokens 118,998 112,458
Output Tokens 30,750 14,344
Reasoning Tokens 491,730 0
Response Time (avg) 48.18s 9.65s
Response Time (max) 422.72s 79.65s
Response Time (total) 1059.93s 212.38s
Parameters 33B total (3B active) 744B total (40B active)
Availability Weights available Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#153 Laguna XS 2.1

medium
Cost
$0.001
Time
30.6s
Tokens
4,678 tok

#138 GLM 5.2

none
Invalid SVG
Cost
$0.033
Time
87.7s
Tokens
7,455 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Laguna XS 2.1 5.5 10.0 33.3% 0 70.35s 7,995 23,767 83,258
GLM 5.2 3.7 9.5 0.0% 0 7.55s 7,263 1,958 0

Quick Compare

Switch Comparison Pair