Navigate
AI BENCHY
Advertise here

Nemotron 3 Super (medium) vs GLM 5.1

The average score is effectively tied at 5.7 vs 5.7. Nemotron 3 Super (medium) has the lower benchmark cost at $0.050 vs $0.164. GLM 5.1 is faster at 6.74s vs 52.31s, with pass rates of 40.9% vs 45.5%.

Last updated at: 2026-09-02

Rank
#212
Total Output Tokens
115,574
Response Time (avg)
52.31s
Total Cost
$0.050
Rank
#216
Total Output Tokens
14,393
Response Time (avg)
6.74s
Total Cost
$0.164
Recommended model GLM 5.1

It has the best score here (5.7), while responding about 7.8x faster than Nemotron 3 Super (medium).

Detailed comparison

Metric Nemotron 3 Super Nemotron 3 Super medium Release: 2026-03-11 Free Available GLM 5.1 GLM 5.1 none Release: 2026-04-07
Score 5.7 5.7
Rank #212 #216
Reliability 9.1 10.0
Consistency 8.9 8.2
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 40.9% 45.5%
Flaky tests 3 5
Total Runs 66 66
Cost per result 0.004 2.073
Total Cost $0.050 $0.164
Input Price $0.085 / 1M $0.966 / 1M
Output Price $0.400 / 1M $3.036 / 1M
Total Input Tokens 81,438 124,219
Output Tokens 17,674 14,393
Reasoning Tokens 97,900 0
Response Time (avg) 52.31s 6.74s
Response Time (max) 431.98s 61.20s
Response Time (total) 1046.21s 148.20s
Parameters 120B total (12B active) 744B total (40B active)
Availability Weights available Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#212 Nemotron 3 Super

medium
Cost
$0.000
Time
272.6s
Tokens
5,296 tok

#216 GLM 5.1

none
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Nemotron 3 Super 3.1 10.0 0.0% 0 147.32s 2,275 797 4,424
GLM 5.1 3.9 9.7 0.0% 0 4.96s 7,256 525 0

Quick Compare

Switch Comparison Pair