Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

GLM 5.3 Flash (max) vs DeepSeek V4 Flash 0731 (high) vs Qwen3.8 27B (high) vs GLM 5.3 (high) benchmark comparison: GLM 5.3 Flash (max) leads on Score with 8.7. GLM 5.3 Flash (max) leads on Reliability with 10.0. DeepSeek V4 Flash 0731 (high) has the lowest Total Cost at $0.058. GLM 5.3 (high) is fastest at 27.82s.

Last updated at: 2026-08-26

Rank
#29
Total Output Tokens
204,340
Response Time (avg)
52.10s
Total Cost
$0.059
Rank
#55
Total Output Tokens
513,307
Response Time (avg)
114.94s
Total Cost
$0.058
Rank
#43
Total Output Tokens
656,635
Response Time (avg)
167.89s
Total Cost
~$0.287
Rank
#100
Total Output Tokens
63,667
Response Time (avg)
27.82s
Total Cost
$0.403
Recommended model GLM 5.3 Flash (max)

It has the best score here (8.7), while costing about 4.3x less than the other models in this comparison.

Detailed comparison

Metric GLM 5.3 Flash GLM 5.3 Flash max Release: 2026-08-26 DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 high Release: 2026-08-01 Qwen3.8 27B Qwen3.8 27B high Release: 2026-08-14 GLM 5.3 GLM 5.3 high Release: 2026-08-20
Score 8.7 8.0 8.4 7.3
Rank #29 #55 #43 #100
Reliability 10.0 10.0 9.6 10.0
Consistency 8.9 8.0 9.2 7.5
Attempts 66/66 66/66 66/66 66/66
Tests Correct
Attempt pass rate 81.8% 75.8% 77.3% 74.2%
Flaky tests 3 5 2 7
Total Runs 66 66 66 66
Cost per result 0.365 0.956 ~1.792 3.099
Total Cost $0.059 $0.058 ~$0.287 $0.403
Input Price $0.075 / 1M $0.060 / 1M N/A $1.400 / 1M
Output Price $0.250 / 1M $0.120 / 1M N/A $4.400 / 1M
Total Input Tokens 96,873 96,702 100,179 87,601
Output Tokens 7,637 16,966 904 6,006
Reasoning Tokens 196,703 496,341 655,731 57,661
Response Time (avg) 52.10s 114.94s 167.89s 27.82s
Response Time (max) 333.47s 600.61s 640.85s 159.91s
Response Time (total) 1146.24s 2528.65s 3693.54s 612.06s
Parameters 320B total (18B active) 284B total (13B active) 27.3B 744B total (40B active)
Availability Open source Closed Weights available Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#29 GLM 5.3 Flash

max
Cost
$0.005
Time
296.5s
Tokens
16,250 tok

#55 DeepSeek V4 Flash 0731

high
Invalid SVG
Cost
$0.000
Time
187.3s
Tokens
19,048 tok

#43 Qwen3.8 27B

high
Cost
~$0.008
Time
270.1s
Tokens
18,963 tok

#100 GLM 5.3

high
Cost
$0.085
Time
362.5s
Tokens
19,310 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GLM 5.3 Flash 10.0 10.0 100.0% 0 43.37s 7,317 381 23,692
DeepSeek V4 Flash 0731 6.3 6.3 55.6% 1 252.67s 7,044 2,160 191,210
Qwen3.8 27B 8.1 7.0 88.9% 1 248.62s 8,235 346 143,335
GLM 5.3 6.6 5.0 66.7% 2 18.89s 7,317 473 8,190

Quick Compare

Switch Comparison Pair