Navigate
Advertise here

GPT-5.2 Chat vs GLM 5.3 FlashX (medium)

GLM 5.3 FlashX (medium) leads on average score with 7.1 vs 7.0. GLM 5.3 FlashX (medium) has the lower benchmark cost at $0.114 vs $0.594. GLM 5.3 FlashX (medium) is faster at 7.37s vs 7.43s, with pass rates of 66.7% vs 69.6%.

Last updated at: 2026-10-01

Compared models

Rank
#153
Total Output Tokens
29,742
Response Time (avg)
7.43s
Total Cost
$0.594
Rank
#150
Total Output Tokens
30,048
Response Time (avg)
7.37s
Total Cost
$0.114
Recommended model GLM 5.3 FlashX (medium)

It has the best score here (7.1), while costing about 5.2x less than GPT-5.2 Chat.

Detailed comparison

Metric GPT-5.2 Chat GPT-5.2 Chat none Release: 2025-12-11 GLM 5.3 FlashX GLM 5.3 FlashX medium Release: 2026-09-21
Score 7.0 7.1
Rank #153 #150
Reliability 10.0 10.0
Consistency 8.7 7.1
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 66.7% 69.6%
Flaky tests 4 8
Total Runs 69 69
Cost per result 4.564 0.946
Total Cost $0.594 $0.114
Input Price $1.750 / 1M $0.370 / 1M
Output Price $14.000 / 1M $1.250 / 1M
Total Input Tokens 101,104 205,100
Output Tokens 29,742 8,639
Reasoning Tokens 0 21,409
Response Time (avg) 7.43s 7.37s
Response Time (max) 38.52s 74.39s
Response Time (total) 163.35s 169.54s
Parameters ~400B total (~17B active) 320B total (18B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#153 GPT-5.2 Chat

none
Cost
$0.010
Time
15.3s
Tokens
797 tok

#150 GLM 5.3 FlashX

medium
Cost
$0.002
Time
8.1s
Tokens
1,480 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.2 Chat 8.8 7.8 88.9% 1 9.82s 7,305 6,731 0
GLM 5.3 FlashX 6.0 4.6 66.7% 2 6.37s 7,317 384 5,562

Quick Compare

Switch Comparison Pair