Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

OpenAI: GPT-5.4 Nano vs Z.ai: GLM 5

Last updated at: 2026-03-17

Metric GPT-5.4 Nano GPT-5.4 Nano medium Release: 2026-03-17 GLM 5 GLM 5 none Release: 2026-02-12
Rank #28 #40
Score 7.4 6.7
Consistency 9.0 10.0
Cost per result 0.769 0.201
Total Cost $0.077 $0.019
Tests Correct
Attempt pass rate 66.7% 52.9%
Flaky tests 2 0
Total Runs 51 51
Output Tokens 2,474 1,551
Reasoning Tokens 54,516 0
Response Time (avg) 11.08s 3.77s
Response Time (max) 94.06s 11.07s
Response Time (total) 188.39s 37.66s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 8.3 10.0 75.0% 0 4.52s 683 2,254
GLM 5 4.8 10.0 25.0% 0 2.37s 275 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 9.8 10.0 100.0% 0 24.13s 349 5,719
GLM 5 3.0 10.0 0.0% 0 4.98s 406 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 10.0 10.0 100.0% 0 2.54s 234 516
GLM 5 10.0 10.0 100.0% 0 5.78s 203 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 5.9 7.2 55.6% 1 38.18s 60 43,325
GLM 5 3.0 10.0 0.0% 0 2.24s 19 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 4.5 10.0 0.0% 0 4.15s 179 443
GLM 5 10.0 10.0 100.0% 0 3.27s 103 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 9.8 10.0 100.0% 0 1.88s 95 521
GLM 5 10.0 10.0 100.0% 0 1.48s 61 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 4.0 7.1 22.2% 1 3.65s 640 1,356
GLM 5 7.7 10.0 66.7% 0 2.05s 264 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 10.0 10.0 100.0% 0 7.71s 234 382
GLM 5 10.0 10.0 100.0% 0 11.07s 220 0

Quick Compare

Switch Comparison Pair