Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

OpenAI: gpt-oss-120b vs Z.ai: GLM 5 Turbo

Last updated at: 2026-03-15

Metric gpt-oss-120b gpt-oss-120b medium Release: 2025-08-05 Free Available GLM 5 Turbo GLM 5 Turbo none Release: 2026-03-15
Rank #44 #53
Score 6.2 5.7
Consistency 7.4 9.5
Cost per result 0.135 0.467
Total Cost $0.010 $0.028
Tests Correct
Attempt pass rate 54.2% 39.6%
Flaky tests 5 1
Total Runs 48 48
Output Tokens 13,210 1,264
Reasoning Tokens 34,230 0
Response Time (avg) 16.65s 2.92s
Response Time (max) 50.92s 8.21s
Response Time (total) 149.88s 46.72s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 7.9 9.8 66.7% 0 19.76s 3,463 2,077
GLM 5 Turbo 3.0 10.0 0.0% 0 3.01s 376 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 10.0 10.0 100.0% 0 31.18s 694 5,072
GLM 5 Turbo 3.0 10.0 0.0% 0 4.89s 144 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 6.4 5.9 66.7% 1 1.98s 241 1,114
GLM 5 Turbo 10.0 10.0 100.0% 0 2.47s 204 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 2.9 4.4 22.2% 2 50.92s 6,784 20,606
GLM 5 Turbo 5.3 10.0 33.3% 0 1.97s 25 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 4.3 10.0 0.0% 0 7.90s 107 387
GLM 5 Turbo 4.2 9.9 0.0% 0 2.18s 48 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 9.9 10.0 100.0% 0 7.63s 126 1,799
GLM 5 Turbo 6.5 10.0 50.0% 0 2.13s 65 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.2 4.7 22.2% 2 11.80s 1,508 2,092
GLM 5 Turbo 5.5 7.4 44.4% 1 2.43s 180 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 9.8 10.0 100.0% 0 6.91s 287 1,083
GLM 5 Turbo 10.0 10.0 100.0% 0 8.21s 222 0

Quick Compare

Switch Comparison Pair